Adaptive distributed new energy power generation equipment cooperative control method and system

Through the adaptive distributed new energy power generation equipment collaborative control method, deep reinforcement learning algorithm and gradient optimization algorithm are used to solve the problem of action conflict and real-time coordination of new energy power generation equipment in the local area of ​​the power grid, efficient coordinated control and conflict suppression are achieved, and grid stability and power generation efficiency are improved.

CN119965994AActive Publication Date: 2025-05-09LONGGANG POWER SUPPLY CO OF STATE GRID ZHEJIANG ELECTRIC POWER CO LTD

Patent Information

Application Number
CN202510436363.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-09
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

In the scenario of abnormal voltage fluctuations in local area of ​​the power grid, the operational conflicts and real-time coordination problems of new energy power generation equipment have not been effectively resolved, making it difficult for the system to quickly and effectively coordinate the movements of each equipment, affecting the overall control effect.

Method used

Adaptive distributed new energy power generation equipment collaborative control method is adopted, and state features are extracted from the device operation parameters collected in real time, action boundaries are determined using deep reinforcement learning algorithms, experience extension data sets are generated, equipment operation reward parameters are calculated, shared experience correction sets are formed, action selection conflict probability is evaluated, and multi-device collaborative control optimization strategy after conflict suppression is iteratively determined.

Benefits of technology

It realizes efficient coordinated control and conflict suppression of new energy power generation equipment in the scenario of abnormal voltage fluctuations, improves the equipment's ability to adapt to abnormal voltage fluctuations, and significantly improves the stability and power generation efficiency of the power grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119965994A_ABST
    Figure CN119965994A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of new energy power generation, in particular to an adaptive distributed new energy power generation equipment cooperative control method and system, and the method comprises the steps: extracting the state feature representation of each new energy power generation equipment from the operation parameters of the new energy power generation equipment; based on the state feature representation, determining an action boundary estimated value of each new energy power generation device by using a deep reinforcement learning algorithm; extracting an equipment state-action data pair according to the action boundary estimation value, and generating an empirical extended data set by simulating different disturbance amplitudes; based on the empirical extended data set, obtaining equipment operation reward parameters of the new energy power generation equipment by using a weighting method; obtaining a sharing experience correction set based on the equipment operation reward parameters; and iteratively determining a multi-device cooperative control optimization strategy by using a gradient optimization algorithm based on the shared experience correction set and the state feature representation. According to the invention, cooperative control of the new energy power generation equipment under abnormal fluctuation of the power grid voltage is realized, and the stability and safety of the power grid are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of renewable energy power generation technology, and in particular to a method and system for collaboratively controlling adaptive distributed renewable energy power generation equipment. Background Art

[0002] In the field of distributed control of renewable energy power generation equipment, especially in the scenario of abnormal voltage fluctuation in local area of ​​power grid, the problems of action conflict and real-time coordination of renewable energy power generation equipment systems are particularly prominent, and there are still a series of technical problems that need to be solved urgently.

[0003] Specifically, when the grid voltage fluctuates rapidly due to a sudden load change, multiple renewable energy power generation devices in the area (such as wind turbines or photovoltaic panels, etc.) need to respond quickly. However, these renewable energy power generation devices have differences in physical properties and operating conditions. For example, some renewable energy power generation devices cannot be further adjusted because their output is close to the upper limit, while other renewable energy power generation devices cannot fully exert their power generation capacity due to factors such as excessive temperature. This inconsistency in state causes the action selection of each device to conflict when facing abnormal voltage fluctuations, making it difficult for the system to quickly and effectively coordinate the actions of each device, which brings great difficulties to collaborative control and affects the overall control effect. In addition, when designing control strategies, how to balance the stability of the grid and the operating cost of equipment is a key problem. If the stability of the grid is overemphasized, some equipment may frequently adjust its output, thereby increasing equipment losses and operating costs; on the contrary, if the operating cost of the equipment is given priority, it may not be able to meet the grid's demand for real-time response, resulting in further deterioration of voltage fluctuations. This contradiction has not been effectively resolved in the existing technology.

[0004] In summary, the existing technology has many problems when dealing with complex scenarios such as abnormal voltage fluctuations in local areas of the power grid. Therefore, it is urgent to seek more effective solutions to improve the adaptability and stability of coordinated control of multiple devices in the power grid and achieve optimal regulation under abnormal voltage fluctuations. Summary of the invention

[0005] In order to solve the above technical problems, the present invention provides a method and system for collaborative control of adaptive distributed renewable energy power generation equipment.

[0006] In a first aspect, the present invention provides a method for cooperatively controlling an adaptive distributed renewable energy power generation device, the method comprising the following steps: Extracting the state characteristic representation of each new energy power generation equipment from the real-time collected operating parameters of the new energy power generation equipment; Based on the state feature representation, a deep reinforcement learning algorithm is used to determine an estimated action boundary value of each new energy power generation device in a voltage abnormal fluctuation scenario; Extracting device state-action data pairs from a historical operation database according to the estimated values ​​of the action boundaries, and generating an empirical extended data set by simulating different disturbance amplitudes according to the device state-action data pairs; Based on the data sampling frequency and the proportion of abnormal scenarios in the empirical extended data set, the equipment operation reward parameters of the new energy power generation equipment are obtained using a weighted method; Based on the device operation reward parameter, extracting state-response data pairs with the same time decay coefficient from a pre-built distributed shared experience pool to form a shared experience correction set; Based on the shared experience correction set and the state feature representation, the action selection conflict probability of each new energy power generation equipment is evaluated, and according to the action selection conflict probability, a gradient optimization algorithm is used to iteratively determine the multi-device collaborative control optimization strategy after conflict suppression.

[0007] In a further embodiment, the step of extracting the state characteristic representation of each new energy power generation equipment from the real-time collected operating parameters of the new energy power generation equipment includes: Collect grid voltage data and operating parameters of new energy power generation equipment in real time, and use a preset data sampling frequency to perform real-time calculations on the grid voltage data to obtain a voltage-frequency deviation value; Comparing the voltage-frequency deviation value with a preset frequency deviation threshold, and if the voltage-frequency deviation value exceeds the preset frequency deviation threshold, extracting a state space description of each new energy power generation device from the operating parameters of the new energy power generation device; Based on the state space description, the principal component analysis method is used to extract the state feature representation of each new energy power generation equipment.

[0008] In a further implementation scheme, the state space description includes the current disturbance simulation amplitude, equipment load factor and output upper limit of the new energy power generation equipment.

[0009] In a further implementation scheme, the step of determining the estimated action boundary value of each new energy power generation device in the voltage abnormal fluctuation scenario using a deep reinforcement learning algorithm based on the state feature representation includes: Introducing action space dimension constraints, using a deep Q network to perform mapping analysis on the state feature representation, and calculating the initial action space range of each new energy power generation device; The initial action space range is detected, and if it is detected that the initial action space range is in a narrow range, a restricted state of the action space in the narrow range is classified to obtain a restricted state classification result; Taking the disturbance simulation amplitude and equipment load factor as clustering features, the K-means clustering algorithm is used to analyze the new energy power generation equipment whose restricted state classification results are restricted, and the restricted level distribution density characteristics are obtained; If the restricted level distribution density feature exceeds a preset distribution density threshold, the restricted level distribution density feature is input into a deep Q network for secondary calculation to generate a dynamic compensation coefficient of the initial action space range; The dynamic compensation coefficient is used to perform dynamic boundary secondary optimization on the initial action space range to obtain an estimated value of the action boundary of each new energy power generation equipment in a voltage abnormal fluctuation scenario.

[0010] In a further embodiment, the step of detecting the initial action space range includes: The initial action space range is segmented using a preset time window length, and the time series variation characteristics of the power regulation range and the temperature limit condition fluctuation data within each time window are extracted; Analyzing the dynamic coupling relationship between the time series variation characteristics of the power adjustment range and the temperature limit condition fluctuation data through a linear regression algorithm, and calculating the dynamic variation trend coefficients of the two; The dynamic change trend coefficient is compared with a preset narrow range threshold, and if the dynamic change trend coefficient is less than the narrow range threshold, it is determined that the initial action space range is in a narrow range.

[0011] In a further embodiment, the step of obtaining the equipment operation reward parameter of the new energy power generation equipment by using a weighted method based on the data sampling frequency and the proportion of abnormal scenarios in the experience extended data set includes: By empirically expanding the data sampling frequency and the proportion of abnormal scenarios in the dataset, the initial reward function value is calculated using the weighted sum method; According to the mapping relationship between the distribution density of the initial reward function value and the proportion of abnormal scenes, a density peak clustering algorithm is used to obtain the reward interval; In each reward interval, the penalty factor benchmark value is calculated according to the load deviation between the equipment load factor and the preset load threshold, and the penalty factor benchmark value is smoothed by the sliding average method to obtain a smoothed dynamic cost penalty distribution; Based on the disturbance simulation amplitude, the dynamic cost penalty distribution in the high disturbance reward interval is compensated and corrected to obtain the corrected penalty distribution; The modified penalty distribution is inversely weighted superimposed on the initial reward function value, and normalized to obtain the equipment operation reward parameter of the new energy power generation equipment.

[0012] In a further embodiment, the step of extracting state-response data pairs having the same time decay coefficient from a pre-built distributed shared experience pool based on the device operation reward parameter to form a shared experience correction set comprises: Extracting a reference time decay coefficient from the device operation reward parameter as a matching reference, and matching the reference time decay coefficient with the time decay coefficient of each data pair in the distributed shared experience pool to extract a state-response data candidate set having the same time decay coefficient; Calculate the response delay difference between the action execution response time of each state-response data pair in the state-response data candidate set and the actual execution timestamp of the device response; Based on the response delay difference, a sliding window average method is used to align the timestamps in the state-response data candidate set to obtain a time-synchronized state-response data set; The state-response data set is divided into different time windows, and the state-response data of multiple renewable energy power generation equipment in the same time window are fused to construct a time-series continuous shared experience correction set.

[0013] In a further embodiment, the step of evaluating the probability of conflict in action selection of each new energy power generation device based on the shared experience correction set and the state feature representation includes: Calculating the abnormal cumulative value of each new energy power generation device according to the state-response data in the shared experience correction set, and identifying high conflict risk intelligent agents from the new energy power generation devices according to the abnormal cumulative value; The voltage abnormal fluctuation amplitude is used as the main sorting basis, and the time decay coefficient is used as the secondary sorting basis. The entropy weight method is used to calculate the action priority of each high conflict risk intelligent agent to obtain the action priority intelligent agent sequence; Based on the action priority agent sequence, the overlapping ratio of power adjustment ranges of high-conflict risk agents with adjacent action priorities is calculated, and the potential conflict boundary is determined according to the overlapping ratio of power adjustment ranges; Map the real-time action instructions to the power adjustment interval of the potential conflict boundary, and count the frequency and magnitude of action instructions crossing the boundary for each high conflict risk agent; The conflict intensity index is calculated according to the product of the action instruction out-of-bounds frequency and the action instruction out-of-bounds amplitude, and the conflict intensity index is normalized to obtain the action selection conflict probability of each new energy power generation equipment.

[0014] In a further embodiment, the step of selecting the conflict probability of the action and iteratively determining the multi-device collaborative control optimization strategy after conflict suppression using a gradient optimization algorithm comprises: According to the conflict probability of action selection of each renewable energy power generation equipment and the real-time state characteristics of the power grid, the corresponding initial disturbance frequency value is determined; Taking minimizing the conflict probability and maximizing the power output efficiency as the optimization goals, based on the optimization goals, using the gradient optimization algorithm to calculate the partial derivatives of the initial disturbance frequency value and the power adjustment range, and obtaining the gradient optimization direction; The disturbance frequency value and power adjustment range are gradually adjusted according to the gradient optimization direction. Through iterative calculation until the optimization target converges, the multi-device collaborative control optimization strategy after conflict suppression is output.

[0015] In a second aspect, the present invention provides an adaptive distributed renewable energy power generation equipment collaborative control system, the system comprising: A data acquisition module is used to extract the status characteristic representation of each new energy power generation equipment from the real-time collected operating parameters of the new energy power generation equipment; A boundary estimation module, used to determine the estimated action boundary value of each new energy power generation equipment in the voltage abnormal fluctuation scenario by using a deep reinforcement learning algorithm based on the state feature representation; A data expansion module, used for extracting device state-action data pairs from a historical operation database according to the action boundary estimation value, and generating an empirical expansion data set by simulating different disturbance amplitudes according to the device state-action data pairs; The reward analysis module is used to obtain the equipment operation reward parameters of the new energy power generation equipment using a weighted method based on the data sampling frequency and the proportion of abnormal scenarios in the experience-expanded data set; An experience correction module, for extracting state-response data pairs with the same time decay coefficient from a pre-built distributed shared experience pool based on the device operation reward parameter to form a shared experience correction set; The strategy generation module is used to evaluate the action selection conflict probability of each new energy power generation equipment based on the shared experience correction set and the state feature representation, and according to the action selection conflict probability, use the gradient optimization algorithm to iteratively determine the multi-device collaborative control optimization strategy after conflict suppression.

[0016] The present invention provides a method and system for cooperative control of adaptive distributed renewable energy power generation equipment, the method comprising extracting a state feature representation of each renewable energy power generation equipment from operating parameters of the renewable energy power generation equipment collected in real time; based on the state feature representation, determining an estimated action boundary of each renewable energy power generation equipment in a voltage abnormal fluctuation scenario using a deep reinforcement learning algorithm; extracting device state-action data pairs from a historical operation database according to the estimated action boundary, and generating an experience extension data set by simulating different disturbance amplitudes based on the device state-action data pairs; obtaining device operation reward parameters of the renewable energy power generation equipment using a weighted method based on the data sampling frequency and the proportion of abnormal scenarios in the experience extension data set; based on the device operation reward parameters, extracting state-response data pairs with the same time attenuation coefficient from a pre-constructed distributed shared experience pool to form a shared experience correction set; based on the shared experience correction set and the state feature representation, evaluating the action selection conflict probability of each renewable energy power generation equipment, and iteratively determining a multi-device cooperative control optimization strategy after conflict suppression using a gradient optimization algorithm based on the action selection conflict probability. Compared with the existing technology, this method adaptively determines the collaborative control strategy of new energy power generation equipment in a complex power grid environment through the combination of deep reinforcement learning algorithm and gradient optimization strategy, realizes efficient collaborative control and conflict suppression of new energy power generation equipment in the scenario of abnormal voltage fluctuation, improves the equipment's adaptability to abnormal voltage fluctuations, and significantly improves the stability and power generation efficiency of the power grid. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic diagram of the process flow of the adaptive distributed renewable energy power generation equipment collaborative control method provided by an embodiment of the present invention; Figure 2 It is a block diagram of an adaptive distributed renewable energy power generation equipment collaborative control system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The following specifically illustrates the implementation mode of the present invention in conjunction with the accompanying drawings. The embodiments are provided for illustrative purposes only and cannot be understood as limiting the present invention. The accompanying drawings are provided for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.

[0019] refer to Figure 1 The embodiment of the present invention provides a method for cooperative control of adaptive distributed renewable energy power generation equipment, such as Figure 1 As shown, the method comprises the following steps: S1. Extract the state characteristic representation of each new energy power generation equipment from the real-time collected operating parameters of the new energy power generation equipment.

[0020] In some embodiments, the step of extracting the state characteristic representation of each new energy power generation device from the real-time collected operating parameters of the new energy power generation device includes: Collect grid voltage data and operating parameters of new energy power generation equipment in real time, and use a preset data sampling frequency to perform real-time calculations on the grid voltage data to obtain a voltage-frequency deviation value; The voltage-frequency deviation value is compared with a preset frequency deviation threshold value. If the voltage-frequency deviation value exceeds the preset frequency deviation threshold value, a state space description of each new energy power generation device is extracted from the operating parameters of the new energy power generation device; the state space description includes the current disturbance simulation amplitude, device load factor and output upper limit of the new energy power generation device; Based on the state space description, the principal component analysis method is used to extract the state feature representation of each new energy power generation equipment.

[0021] Specifically, this embodiment collects grid voltage data and operating parameters of each renewable energy power generation equipment in real time through a sensor network installed in a renewable energy power station, and uses fast Fourier transform to calculate the grid voltage data in real time according to a preset data sampling frequency to obtain the voltage frequency value of each sampling point, and then compares the voltage frequency value of each sampling point with the standard frequency to obtain the voltage frequency deviation value. These voltage frequency deviation values ​​reflect the real-time fluctuation of the grid voltage frequency. This embodiment compares each voltage frequency deviation value with a preset frequency deviation threshold to determine whether the grid voltage frequency is within a normal range. If the voltage frequency deviation value exceeds the preset frequency deviation threshold, it indicates that the grid voltage frequency is abnormal and further analysis of the state of the renewable energy power generation equipment is required. At this time, this embodiment extracts key state parameters such as the current disturbance simulation amplitude, equipment load factor and output upper limit for each renewable energy power generation equipment to obtain a state space description of each renewable energy power generation equipment. The description includes the key operating characteristics of the equipment under the current power grid conditions, among which the disturbance simulation amplitude reflects the degree of disturbance suffered by the equipment under abnormal voltage conditions, which can be estimated by comparing parameters such as the output power change of the equipment before and after the abnormality in this embodiment; the equipment load factor represents the proportion of the current load of the equipment to its maximum design load, which can be calculated by the ratio of real-time output power to rated power in this embodiment; the output upper limit represents the maximum output power that the equipment can reach under the current voltage conditions; then this embodiment organizes the extracted state space descriptions of each new energy power generation equipment into a matrix form, where each row represents the state space description of a device at a certain moment, and uses principal component analysis (PCA) to process the data in the state space description, extracts the principal components that have the greatest impact on the equipment operating status, and uses the principal components extracted by PCA as the state feature representations of each new energy power generation equipment, which can effectively reflect the operating status of the equipment under abnormal voltage conditions.

[0022] S2. Based on the state feature representation, a deep reinforcement learning algorithm is used to determine the estimated action boundary value of each new energy power generation equipment in the scenario of abnormal voltage fluctuation.

[0023] In some embodiments, the step of determining the estimated action boundary value of each new energy power generation device in the abnormal voltage fluctuation scenario using a deep reinforcement learning algorithm based on the state feature representation includes: Introducing action space dimension constraints, using a deep Q network to perform mapping analysis on the state feature representation, and calculating the initial action space range of each new energy power generation device; The initial action space range is detected, and if it is detected that the initial action space range is in a narrow range, a restricted state of the action space in the narrow range is classified to obtain a restricted state classification result; Taking the disturbance simulation amplitude and equipment load factor as clustering features, the K-means clustering algorithm is used to analyze the new energy power generation equipment whose restricted state classification results are restricted, and the restricted level distribution density characteristics are obtained; If the restricted level distribution density feature exceeds a preset distribution density threshold, the restricted level distribution density feature is input into a deep Q network for secondary calculation to generate a dynamic compensation coefficient of the initial action space range; The dynamic compensation coefficient is used to perform dynamic boundary secondary optimization on the initial action space range to obtain an estimated value of the action boundary of each new energy power generation equipment in a voltage abnormal fluctuation scenario.

[0024] Specifically, this implementation first determines the action type of the new energy power generation equipment in the voltage abnormal fluctuation scenario according to the type, performance parameters and grid access requirements of the new energy power generation equipment, such as power regulation, equipment protection, etc., and defines the action space dimension according to the action type. For example, the action space dimension of the photovoltaic inverter may include the output power adjustment amplitude, the power factor adjustment range, etc.; the wind turbine generator set may include the pitch angle adjustment range, the generator torque adjustment amplitude, etc. This embodiment takes into account the physical characteristics and operation safety of the new energy power generation equipment, and constrains the dimension of the action space to ensure that all possible actions meet the safety operation standards of the equipment. For example, the output power adjustment amplitude of the photovoltaic inverter is limited to within ±20% of its rated power, and then constructs a deep Q network (DQN), which consists of a multi-layer neural network The network is composed of a plurality of input layer neurons, the number of neurons in the input layer matches the dimension of the state feature representation, the hidden layer adopts a multi-layer perceptron structure, the activation function selects ReLU, and the number of neurons in the output layer corresponds to the dimension of the action space of the new energy power generation equipment. Based on the action space dimension constraint, this embodiment inputs the historical state feature representation data into the deep Q network for training, so that it can accurately map the initial action space range of each new energy power generation equipment according to the input state feature representation, and obtain the initial action space range of each new energy power generation equipment through the forward propagation calculation of the network. The initial action space range includes the upper limit and lower limit of the power adjustment range and the threshold interval of the temperature limit condition, and the initial action space range is detected. In some embodiments, the step of detecting the initial action space range includes: The initial action space range is segmented using a preset time window length, and the time series variation characteristics of the power regulation range and the temperature limit condition fluctuation data within each time window are extracted; Analyzing the dynamic coupling relationship between the time series variation characteristics of the power adjustment range and the temperature limit condition fluctuation data through a linear regression algorithm, and calculating the dynamic variation trend coefficients of the two; The dynamic change trend coefficient is compared with a preset narrow range threshold, and if the dynamic change trend coefficient is less than the narrow range threshold, it is determined that the initial action space range is in a narrow range.

[0025] Specifically, this embodiment sets the time window length according to the duration characteristics of abnormal fluctuations in grid voltage and the response speed requirements of new energy power generation equipment, divides the initial action space range into segments according to the time window length, and obtains action space sub-ranges within multiple time windows. For each time window, the time series change characteristics of the power regulation range of the new energy power generation equipment are extracted. Taking the wind turbine generator set as an example, this embodiment can calculate the characteristic values ​​such as the change rate and change trend of the pitch angle adjustment amplitude and the generator torque adjustment amplitude within each time window, and extract the temperature limit condition fluctuation data at the same time. The temperature limit condition fluctuation data can specifically include the ambient temperature change range, key components inside the equipment (such as inverter power module, generator stator winding Then, this embodiment uses the time series change characteristics of the power regulation range and the temperature limit condition fluctuation data as variables to construct a linear regression model, and calculates the regression coefficient between the two by the least squares method and other methods to obtain the dynamic change trend coefficient of the two. The dynamic change trend coefficient reflects the synchronization or correlation of the two over time. At the same time, this embodiment sets a narrow range threshold according to the operating experience of new energy power generation equipment and the stability requirements of power grid operation. The narrow range threshold is used to determine whether the initial action space range is in the narrow range. If the dynamic change trend coefficient is less than the narrow range threshold, it is determined that the initial action space range is in the narrow range. At this time, the action selection of the device is greatly restricted.

[0026] Then, this embodiment uses a support vector machine to classify the restricted state of the action space in a narrow range. For example, the restricted state is divided into three categories: mild restriction, moderate restriction and severe restriction, and the restricted state classification result is obtained. Then, the disturbance simulation amplitude and the equipment load coefficient are used as clustering features to initialize the K-means clustering algorithm. Through iterative calculation, each cluster center and the sample belonging category are obtained, and then the restricted level distribution density feature is obtained. The restricted level distribution density feature reflects the distribution of devices at different restricted levels. If the restricted level distribution density feature exceeds the preset distribution density threshold, the restricted level distribution density feature is input into the deep Q network for secondary calculation. This embodiment adds a compensation coefficient calculation module in the deep Q network. The compensation coefficient calculation module takes the restricted level distribution density feature as input and generates an initial The dynamic compensation coefficient of the initial action space range is used to adjust the initial action space range to cope with the limitations brought by the narrow range. For example, for a certain new energy power generation equipment, the generated dynamic compensation coefficient is 1.2, indicating that the initial action space range needs to be appropriately expanded and compensated. Finally, this embodiment performs element-wise multiplication of the dynamic compensation coefficient and the initial action space range to obtain an optimized action boundary estimate. For example, the initial output power adjustment amplitude range is [-15%, +18%], and the dynamic compensation coefficient is 1.2, then the optimized output power adjustment amplitude range is [-18%, +21.6%]. The optimized action boundary estimate can more accurately reflect the actual action capability of the new energy power generation equipment in the voltage abnormal fluctuation scenario, and provide strong support for the intelligent control of the equipment.

[0027] S3. According to the action boundary estimation value, extract the device state-action data pair from the historical operation database, and generate an empirical extended data set by simulating different disturbance amplitudes according to the device state-action data pair.

[0028] Specifically, this embodiment selects device state data and corresponding action data related to action boundary estimation values ​​from a historical operation database to form device state-action data pairs, which contain operating parameters of the device in different states and corresponding operation records. Then, this embodiment performs preprocessing operations on the extracted device state-action data pairs, and the preprocessing operations include but are not limited to data cleaning and data standardization to ensure data quality. Then, feature dimensions that can characterize the diversity of device operating states are selected from the extracted device state-action data pairs to construct cluster feature vectors, such as grid voltage and frequency deviation values, equipment load factors, disturbance simulation amplitudes, etc. The cluster feature vectors are input into a K-means clustering algorithm for cluster analysis, and the clustering categories of each data pair are obtained through iterative calculations, thereby obtaining classified state distribution characteristics, and then calculating the state coverage breadth represented by each cluster center based on the classified state distribution characteristics, wherein the state coverage breadth can be measured by the number or density of data points around the cluster center.

[0029] This embodiment divides the device state data into different categories according to the state coverage breadth, and extracts the state distribution features of each category. These features may include the coordinates of the cluster center, the cluster radius, the state density, etc. For each classified state distribution feature, it is determined whether its state coverage breadth is lower than the preset coverage breadth threshold. If the state coverage breadth is lower than the coverage breadth threshold, it is considered that the state data of this category is insufficient and needs to be empirically expanded. A seed data pair is randomly selected from the category that needs to be empirically expanded as a basis. At the same time, the type and amplitude distribution of the disturbance suffered by the equipment in the historical operation data are analyzed to determine the range of the disturbance simulation amplitude. and step size. On the basis of the original device state-action data pair, according to different ranges and step sizes of disturbance simulation amplitudes, the device state feature representation in the seed data pair is disturbed and simulated to generate a new device state feature representation, and the action data corresponding to the new device state feature representation is simulated. For example, according to the MPPT (maximum power point tracking) algorithm and inverter control strategy of the photovoltaic inverter, the changes in the output power regulation amplitude and the power factor adjustment value under the input power after the disturbance are calculated, so as to obtain the action data corresponding to the new state, and the generated new state-action data pair is added to the original data set to form an empirical extended data set.

[0030] S4. Based on the data sampling frequency and the proportion of abnormal scenarios in the empirical extended data set, the equipment operation reward parameters of the new energy power generation equipment are obtained using a weighted method.

[0031] In some embodiments, the step of obtaining the equipment operation reward parameter of the new energy power generation equipment by using a weighted method based on the data sampling frequency and the abnormal scene proportion in the experience extended data set includes: By empirically expanding the data sampling frequency and the proportion of abnormal scenarios in the dataset, the initial reward function value is calculated using the weighted sum method; According to the mapping relationship between the distribution density of the initial reward function value and the proportion of abnormal scenes, a density peak clustering algorithm is used to obtain the reward interval; In each reward interval, the penalty factor benchmark value is calculated according to the load deviation between the equipment load factor and the preset load threshold, and the penalty factor benchmark value is smoothed by the sliding average method to obtain a smoothed dynamic cost penalty distribution; Based on the disturbance simulation amplitude, the dynamic cost penalty distribution in the high disturbance reward interval is compensated and corrected to obtain the corrected penalty distribution; The modified penalty distribution is inversely weighted superimposed on the initial reward function value, and normalized to obtain the equipment operation reward parameter of the new energy power generation equipment.

[0032] Specifically, this embodiment extracts the data sampling frequency values ​​of all sampling moments from the experience extended data set, finds the maximum frequency value therein, and normalizes the data sampling frequency in the experience extended data set according to the maximum frequency value, that is, calculates the ratio of the frequency at each sampling moment to the maximum frequency value to obtain the frequency weight coefficient, and at the same time counts the number of data points of various abnormal scenarios (such as voltage sag or frequency fluctuation, etc.) in the experience extended data set, and calculates the proportion of each type of abnormal scenario according to the ratio between the number of data points of each type of abnormal scenario and the total number of data points, and according to the proportion of abnormal scenarios in the experience extended data set, an inverse proportional function is used to normalize the data points of various types of abnormal scenarios. Assign scenario weight coefficients, the higher the proportion of abnormal scenarios, the lower the corresponding scenario weight coefficient. Then, this embodiment extracts the power output efficiency and voltage deviation value from the operating parameters of the new energy power generation equipment, and weightedly sums the power output efficiency and the voltage deviation value based on the frequency weight coefficient and the scenario weight coefficient. Specifically, for each sampling moment, this embodiment multiplies the power output efficiency by the frequency weight coefficient to obtain the power output efficiency weight value, and multiplies the voltage deviation value by the scenario weight coefficient to obtain the voltage deviation weight value. Then, the power output efficiency weight value and the voltage deviation weight value are weightedly summed to obtain the initial reward function value at each sampling moment.

[0033] Next, this embodiment divides the initial reward function value according to the preset time window to obtain the reward function value sequence in multiple time windows, and performs statistical analysis on the initial reward function value in each time window to obtain the reward function value distribution density. For example, the histogram method is used to count the proportion of data points in each reward value interval. According to the mapping relationship between the reward function value distribution density and the abnormal scene proportion, the reward function value distribution density is used as the basis for selecting the cluster center. Combined with the gradient change of the abnormal scene proportion, the density peak clustering algorithm is used to divide the initial reward function value into several reward function intervals. Specifically, this embodiment uses the reward function value distribution density as the basis for selecting the cluster center. The local density of each reward function value point and the distance to the nearest high-density point are calculated in combination with the gradient change of the abnormal scene proportion; then, the point with high local density and long distance is selected as the cluster center; finally, the other points are assigned to the nearest cluster center to form a high reward interval, a medium reward interval and a low reward interval, and a priority label is assigned to each interval. The high reward interval corresponds to a low abnormal scene proportion and has the highest priority. In each reward function interval, the penalty factor baseline value is calculated according to the load deviation between the device load coefficient and the preset load threshold (the absolute difference between the actual load and the threshold as a percentage of the threshold). The greater the deviation, the higher the penalty factor. The higher the penalty value, the smoothing process is performed on the penalty factor benchmark value by the sliding average method. At each data point, the penalty factor benchmark values ​​of the two data points before and after are averaged to eliminate the instantaneous load fluctuation noise, and the smoothed dynamic cost penalty distribution is obtained. The voltage fluctuation weight coefficient is multiplied and fused with the disturbance simulation amplitude to generate a boundary adjustment coefficient. In this embodiment, the boundary adjustment coefficient is equal to the voltage fluctuation weight multiplied by the normalized value of the disturbance amplitude. The normalized value of the disturbance amplitude is the ratio of the current disturbance amplitude to the maximum disturbance amplitude. The high priority reward interval in the dynamic cost penalty distribution is compensated and corrected according to the boundary adjustment coefficient. The compensation amount It is equal to the penalty value multiplied by the boundary adjustment coefficient. In this step, by integrating the voltage fluctuation and disturbance amplitude information, the reward parameter deviation in the high disturbance scenario is compensated first, and the corrected penalty distribution is inversely weighted superimposed with the initial reward function value. Specifically, the reverse superposition is performed according to the priority weight, with the high priority weight accounting for 70%, and the medium and low priorities each accounting for 15%. Through the range normalization processing, the superimposed result is mapped to the [0, 1] interval to obtain the equipment operation reward parameters of the new energy power generation equipment. These equipment operation reward parameters can comprehensively reflect the performance of the equipment under different operating conditions, and provide strong support for the optimized operation and scheduling of the equipment.

[0034] S5. Based on the device running reward parameters, extract state-response data pairs with the same time decay coefficient from a pre-built distributed shared experience pool to form a shared experience correction set.

[0035] In some embodiments, the step of extracting state-response data pairs having the same time decay coefficient from a pre-built distributed shared experience pool based on the device operation reward parameter to form a shared experience correction set includes: Extracting a reference time decay coefficient from the device operation reward parameter as a matching reference, and matching the reference time decay coefficient with the time decay coefficient of each data pair in the distributed shared experience pool to extract a state-response data candidate set having the same time decay coefficient; Calculate the response delay difference between the action execution response time of each state-response data pair in the state-response data candidate set and the actual execution timestamp of the device response; Based on the response delay difference, a sliding window average method is used to align the timestamps in the state-response data candidate set to obtain a time-synchronized state-response data set; The state-response data set is divided into different time windows, and the state-response data of multiple renewable energy power generation equipment in the same time window are fused to construct a time-series continuous shared experience correction set.

[0036] Specifically, this embodiment extracts a reference time decay coefficient from the device operation reward parameter, and the reference time decay reflects the decay characteristics of the reward parameter over time. For example, in the reward parameter of the photovoltaic inverter, it is assumed that its reference time decay coefficient is 0.95, which means that after each unit time, the reward value decays to 95% of the original value. The pre-constructed distributed shared experience pool is traversed. The pre-constructed distributed shared experience pool stores a large number of state-response data pairs from different new energy power generation equipment. Each data pair contains information such as device state characteristics, action execution records, response timestamps, and corresponding time decay coefficients. For each data pair in the experience pool, the absolute difference percentage between its time decay coefficient and the reference time decay coefficient is calculated. Then, this embodiment sets a preset absolute difference threshold. If the difference percentage is less than or equal to the absolute difference threshold, it is determined that the time decay coefficient of the data pair is consistent with the reference time decay coefficient. All qualified data pairs are screened out to form a state-response data candidate set, and invalid data with excessive time decay differences are eliminated.

[0037] For each data pair in the state-response data candidate set, the difference between the response time after the action is executed and the actual execution timestamp of the device response is calculated to obtain the response delay difference. At the same time, this embodiment sets a dynamic delay threshold. If the response delay difference exceeds the dynamic delay threshold, the data pair whose response delay difference exceeds the dynamic delay threshold is determined as misleading data caused by the device response delay and is removed. For the remaining data, this embodiment uses a sliding window average method to align the timestamps. The sliding window length can be set to twice the delay difference mean. The timestamps of each data pair are averaged within the window to obtain a new aligned timestamp, thereby eliminating the clock deviation of the distributed device. After the alignment process, a time-synchronized state-response data set is obtained. Then, this embodiment The time-synchronized state-response data set is divided into preset time windows, and the state-response data of multiple devices in the same time window are fused using the Kalman filter algorithm. Specifically, a state space model is constructed with voltage and power as state variables. For example, the state variables of a photovoltaic inverter may include DC side voltage, AC side output power, etc. The state value at the current moment is then predicted based on the state estimate at the previous moment. The predicted value is corrected using the multi-device data actually observed at the current moment, and the state estimate and covariance matrix are updated to fuse the multi-device data and suppress noise interference. A time continuity label is added to the fused data to ensure that there are no breakpoints in the data on the time axis. Finally, a time-series continuous and noise-suppressed multi-device fused data set is obtained, thereby forming a shared experience correction set.

[0038] S6. Based on the shared experience correction set and the state feature representation, the action selection conflict probability of each new energy power generation equipment is evaluated, and according to the action selection conflict probability, a gradient optimization algorithm is used to iteratively determine the multi-device collaborative control optimization strategy after conflict suppression.

[0039] In some embodiments, the step of evaluating the probability of conflict in action selection of each new energy power generation device based on the shared experience correction set and the state feature representation includes: Calculating the abnormal cumulative value of each new energy power generation device according to the state-response data in the shared experience correction set, and identifying high conflict risk intelligent agents from the new energy power generation devices according to the abnormal cumulative value; The voltage abnormal fluctuation amplitude is used as the main sorting basis, and the time decay coefficient is used as the secondary sorting basis. The entropy weight method is used to calculate the action priority of each high conflict risk intelligent agent to obtain the action priority intelligent agent sequence; Based on the action priority agent sequence, the overlapping ratio of power adjustment ranges of high-conflict risk agents with adjacent action priorities is calculated, and the potential conflict boundary is determined according to the overlapping ratio of power adjustment ranges; Map the real-time action instructions to the power adjustment interval of the potential conflict boundary, and count the frequency and magnitude of action instructions crossing the boundary for each high conflict risk agent; The conflict intensity index is calculated according to the product of the action instruction out-of-bounds frequency and the action instruction out-of-bounds amplitude, and the conflict intensity index is normalized to obtain the action selection conflict probability of each new energy power generation equipment.

[0040] Specifically, this embodiment extracts state-response data from the shared experience correction set. These data include the real-time state (such as voltage, current, power, etc.) and corresponding responses (such as adjustment actions, fault information, etc.) of the new energy power generation equipment. According to the state-response data of each new energy power generation equipment, its abnormal cumulative value within a period of time is calculated. Specifically, it can be calculated by counting the number of times the abnormal state occurs, and the new energy power generation equipment whose abnormal cumulative value exceeds the abnormal threshold is identified as a high-conflict risk intelligent body. Then, this embodiment extracts the voltage abnormal fluctuation amplitude data of each high-conflict risk intelligent body from the shared experience correction set as the main sorting basis. The voltage abnormal fluctuation amplitude can be expressed as The absolute value of the difference between the actual voltage and the rated voltage reflects the instability of the voltage state of the equipment. The time decay coefficient of each high-conflict risk agent is extracted as the secondary sorting basis. The time decay coefficient is used to consider the impact of historical anomalies on the current state. Normalization is performed according to the voltage abnormality fluctuation amplitude and the time decay coefficient to eliminate dimension and magnitude differences. The entropy weight method is used to calculate the action priority of each high-conflict risk agent. The entropy weight method determines its weight by calculating the entropy value of each sorting basis. The comprehensive sorting score of each high-conflict risk agent is calculated according to the weight. The high-conflict risk agents are sorted from high to low according to the comprehensive sorting score to obtain the action priority agent sequence.

[0041] In this embodiment, in the action priority agent sequence, the power adjustment range data of the high-conflict risk agents with adjacent action priorities are obtained, the power adjustment range represents the output power interval that the device can adjust under the current operating state, the length of the overlapping portion of the power adjustment range of the high-conflict risk agents with adjacent action priorities is calculated, the overlapping ratio of the power adjustment range of the high-conflict risk agents with adjacent action priorities is calculated according to the length of the overlapping portion of the power adjustment range, the area where the power adjustment range overlapping ratio exceeds the overlapping threshold is determined as the potential conflict boundary, then the real-time action instructions of each high-conflict risk agent are obtained, and the real-time action instructions are mapped to the power adjustment interval of the potential conflict boundary to evaluate whether the action instructions are out of bounds. For example, if the power adjustment range of the potential conflict boundary is [150, 250], and the real-time action of a certain agent is The power adjustment value corresponding to the instruction is 220, which falls within the conflict boundary. At the same time, the frequency and amplitude of action instruction crossing the boundary of each high conflict risk agent are counted within a certain time range. The frequency of crossing the boundary refers to the number of times the action instruction exceeds its own power adjustment range or the power adjustment interval of the potential conflict boundary. The amplitude of crossing the boundary refers to the absolute value of the difference between the crossing action instruction value and the boundary value of the power adjustment range. According to the product of the frequency of action instruction crossing the boundary and the amplitude of action instruction crossing the boundary, the conflict intensity index of each high conflict risk agent is calculated. This conflict intensity index reflects the severity of the action instruction crossing the boundary. The conflict intensity indexes of all high conflict risk agents are normalized to obtain the action selection conflict probability of each new energy power generation equipment. The normalization process can ensure that the conflict probability values ​​of all equipment are between 0 and 1, which is convenient for comparison and evaluation.

[0042] In some implementations, the step of selecting the conflict probability according to the action and iteratively determining the multi-device collaborative control optimization strategy after conflict suppression using a gradient optimization algorithm includes: According to the conflict probability of action selection of each renewable energy power generation equipment and the real-time state characteristics of the power grid, the corresponding initial disturbance frequency value is determined; Taking minimizing the conflict probability and maximizing the power output efficiency as the optimization goals, based on the optimization goals, using the gradient optimization algorithm to calculate the partial derivatives of the initial disturbance frequency value and the power adjustment range, and obtaining the gradient optimization direction; The disturbance frequency value and power adjustment range are gradually adjusted according to the gradient optimization direction. Through iterative calculation until the optimization target converges, the multi-device collaborative control optimization strategy after conflict suppression is output.

[0043] Specifically, this embodiment collects the action selection conflict probability and the real-time state characteristic data of each new energy power generation equipment, wherein the real-time state characteristics of the power grid may include key parameters such as the power grid voltage deviation rate and the load rate. This embodiment queries the pre-constructed conflict-disturbance mapping relationship table according to the action selection conflict probability, obtains the basic disturbance frequency value, and dynamically compensates the basic disturbance frequency value according to the real-time state characteristics of the power grid, and adds the basic disturbance frequency value to the dynamic compensation value to obtain the dynamically adjusted initial disturbance frequency value. For example, when the voltage deviation rate increases by 1%, the disturbance frequency value is compensated by 0.5Hz; when the load rate increases by 10%, the disturbance frequency value is compensated by 0.3Hz. Then, this embodiment defines the optimization goal as minimizing the action selection conflict probability and maximizing the power output efficiency, so as to ensure that the power output of the equipment is increased as much as possible while reducing the conflict, and construct a multi-objective optimization function, which is specifically: In the formula, Optimize function values ​​for multiple objectives; Select conflict probabilities for actions; is the power output efficiency; and They are all weight coefficients. The weight coefficients can be dynamically adjusted according to the real-time stability of the grid voltage to balance the relationship between conflict probability suppression and power output efficiency improvement.

[0044] This embodiment uses a numerical differentiation method to calculate the partial derivatives of the multi-objective optimization function value with respect to the disturbance frequency value and the power adjustment range to obtain the rate of change of the objective function value. The direction in which the objective function value changes the fastest is found according to the rate of change of the objective function value, and the direction of gradient descent is determined. For example, if the calculation result of the partial derivative of the initial disturbance frequency value is negative, it indicates that the initial disturbance frequency value is increased, and vice versa. Similarly, the adjustment direction of the power adjustment range is determined. This embodiment gradually adjusts the disturbance frequency value and the power adjustment range in the direction of gradient optimization. The new initial disturbance frequency value is the difference between the initial disturbance frequency value and the partial derivative of the initial disturbance frequency value, and the new power adjustment range is the power adjustment range. The difference between the initial disturbance frequency value and the partial derivative of the power regulation range is calculated. After each adjustment, the multi-objective optimization function value is recalculated according to the updated initial disturbance frequency value and the power regulation range. When the change rate of the multi-objective optimization function value for three consecutive iterations is less than the preset change rate threshold, the optimization process is judged to have converged and the iterative calculation is terminated. When the optimization target converges, the multi-device collaborative control optimization strategy after conflict suppression is output. This strategy includes the final disturbance frequency value and power regulation range of each device, as well as the corresponding control instructions. The optimization strategy is deployed in the actual new energy power generation equipment control system to guide the operation and control of the equipment, and to achieve effective conflict suppression and efficient operation of the system.

[0045] An embodiment of the present invention provides an adaptive distributed renewable energy power generation equipment collaborative control method, the method comprising extracting a state feature representation of each renewable energy power generation equipment from operating parameters of the renewable energy power generation equipment collected in real time; based on the state feature representation, determining an action boundary estimation value of each renewable energy power generation equipment in a voltage abnormal fluctuation scenario using a deep reinforcement learning algorithm; extracting a device state-action data pair from a historical operation database according to the action boundary estimation value, and generating an experience extension data set by simulating different disturbance amplitudes based on the device state-action data pair; obtaining a device operation reward parameter of the renewable energy power generation equipment using a weighted method based on the data sampling frequency and the proportion of abnormal scenarios in the experience extension data set; extracting a state-response data pair with the same time attenuation coefficient from a pre-constructed distributed shared experience pool based on the device operation reward parameter to form a shared experience correction set; based on the shared experience correction set and the state feature representation, evaluating the action selection conflict probability of each renewable energy power generation equipment, and iteratively determining a multi-device collaborative control optimization strategy after conflict suppression using a gradient optimization algorithm based on the action selection conflict probability. Compared with the existing technology, this method adaptively determines the collaborative control strategy of new energy power generation equipment in a complex power grid environment through the combination of deep reinforcement learning algorithm and gradient optimization strategy, realizes efficient collaborative control and conflict suppression of new energy power generation equipment in the scenario of abnormal voltage fluctuation, improves the equipment's adaptability to abnormal voltage fluctuations, and significantly improves the stability and power generation efficiency of the power grid.

[0046] It should be noted that the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0047] In one embodiment, Figure 2 As shown, an embodiment of the present invention provides an adaptive distributed new energy power generation equipment collaborative control system, the system comprising: The data acquisition module 101 is used to extract the state characteristic representation of each new energy power generation device from the real-time collected operating parameters of the new energy power generation device; A boundary estimation module 102 is used to determine the estimated action boundary value of each new energy power generation device in the voltage abnormal fluctuation scenario based on the state feature representation using a deep reinforcement learning algorithm; A data expansion module 103 is used to extract device state-action data pairs from a historical operation database according to the action boundary estimation value, and generate an empirical expansion data set by simulating different disturbance amplitudes according to the device state-action data pairs; The reward analysis module 104 is used to obtain the equipment operation reward parameters of the new energy power generation equipment by using a weighted method based on the data sampling frequency and the abnormal scene proportion in the experience extended data set; An experience correction module 105 is used to extract state-response data pairs with the same time decay coefficient from a pre-built distributed shared experience pool based on the device operation reward parameter to form a shared experience correction set; The strategy generation module 106 is used to evaluate the action selection conflict probability of each new energy power generation device based on the shared experience correction set and the state feature representation, and use the gradient optimization algorithm to iteratively determine the multi-device collaborative control optimization strategy after conflict suppression based on the action selection conflict probability.

[0048] For the specific definition of an adaptive distributed renewable energy power generation equipment collaborative control system, please refer to the above-mentioned definition of an adaptive distributed renewable energy power generation equipment collaborative control method, which will not be repeated here. A person of ordinary skill in the art will appreciate that the various modules and steps described in conjunction with the embodiments disclosed in this application can be implemented in hardware, software, or a combination of both. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0049] The embodiment of the present invention provides an adaptive distributed new energy power generation equipment collaborative control system, wherein the system extracts state feature representations of each new energy power generation equipment from real-time collected operation parameters of the new energy power generation equipment through a data acquisition module; a boundary estimation module determines the action boundary estimation value of each new energy power generation equipment in the voltage abnormal fluctuation scenario by using a deep reinforcement learning algorithm based on the state feature representation; a data extension module extracts device state-action data pairs from a historical operation database according to the action boundary estimation value, and generates an experience extension data set by simulating different disturbance amplitudes according to the device state-action data pairs; a reward analysis module obtains device operation reward parameters of the new energy power generation equipment by using a weighted method based on the data sampling frequency and the abnormal scenario proportion in the experience extension data set; an experience correction module extracts state-response data pairs with the same time attenuation coefficient from a pre-constructed distributed shared experience pool based on the device operation reward parameters to form a shared experience correction set; a strategy generation module evaluates the action selection conflict probability of each new energy power generation equipment based on the shared experience correction set and the state feature representation, and iteratively determines the multi-device collaborative control optimization strategy after conflict suppression by using a gradient optimization algorithm according to the action selection conflict probability. Compared with the existing technology, this system adaptively determines the collaborative control strategy of new energy power generation equipment in a complex power grid environment through the combination of deep reinforcement learning algorithm and gradient optimization strategy, realizes efficient collaborative control and conflict suppression of new energy power generation equipment in the scenario of abnormal voltage fluctuation, improves the equipment's adaptability to abnormal voltage fluctuations, and significantly improves the stability and power generation efficiency of the power grid.

[0050] The above-mentioned embodiments only express several preferred implementation modes of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in the technical field, several improvements and substitutions can be made without departing from the technical principles of the present invention, and these improvements and substitutions should also be regarded as the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be based on the protection scope of the claims.

Claims

1. An adaptive distributed renewable energy power generation equipment collaborative control method, characterized in that: The following steps are involved: Extracting the state characteristic representation of each new energy power generation equipment from the real-time collected operating parameters of the new energy power generation equipment; Based on the state feature representation, a deep reinforcement learning algorithm is used to determine an estimated action boundary value of each renewable energy power generation device in a voltage abnormal fluctuation scenario; Extracting device state-action data pairs from a historical operation database according to the action boundary estimation value, and generating an empirical extended data set by simulating different disturbance amplitudes according to the device state-action data pairs; Based on the data sampling frequency and the proportion of abnormal scenarios in the empirical extended data set, the equipment operation reward parameters of the new energy power generation equipment are obtained using a weighted method; Based on the device operation reward parameter, extracting state-response data pairs with the same time decay coefficient from a pre-built distributed shared experience pool to form a shared experience correction set; Based on the shared experience correction set and the state feature representation, the action selection conflict probability of each new energy power generation equipment is evaluated, and according to the action selection conflict probability, a gradient optimization algorithm is used to iteratively determine the multi-device collaborative control optimization strategy after conflict suppression.

2. The adaptive distributed renewable energy power generation equipment collaborative control method according to claim 1, characterized in that: The step of extracting the state characteristic representation of each new energy power generation equipment from the real-time collected operating parameters of the new energy power generation equipment includes: Collect grid voltage data and operating parameters of new energy power generation equipment in real time, and use a preset data sampling frequency to perform real-time calculations on the grid voltage data to obtain a voltage-frequency deviation value; Comparing the voltage-frequency deviation value with a preset frequency deviation threshold, and if the voltage-frequency deviation value exceeds the preset frequency deviation threshold, extracting a state space description of each new energy power generation device from the operating parameters of the new energy power generation device; Based on the state space description, the principal component analysis method is used to extract the state feature representation of each new energy power generation equipment.

3. The adaptive distributed renewable energy power generation equipment collaborative control method according to claim 2, characterized in that: The state space description includes the current disturbance simulation amplitude, equipment load factor and output upper limit of the new energy power generation equipment.

4. The adaptive distributed renewable energy power generation equipment collaborative control method according to claim 1, characterized in that: The step of determining the estimated action boundary value of each new energy power generation device in the voltage abnormal fluctuation scenario by using a deep reinforcement learning algorithm based on the state feature representation includes: Introducing action space dimension constraints, using a deep Q network to perform mapping analysis on the state feature representation, and calculating the initial action space range of each new energy power generation device; The initial action space range is detected, and if it is detected that the initial action space range is in a narrow range, a restricted state of the action space in the narrow range is classified to obtain a restricted state classification result; Taking the disturbance simulation amplitude and equipment load factor as clustering features, the K-means clustering algorithm is used to analyze the new energy power generation equipment whose restricted state classification results are restricted, and the restricted level distribution density characteristics are obtained; If the restricted level distribution density feature exceeds a preset distribution density threshold, the restricted level distribution density feature is input into a deep Q network for secondary calculation to generate a dynamic compensation coefficient of the initial action space range; The dynamic compensation coefficient is used to perform dynamic boundary secondary optimization on the initial action space range to obtain an estimated value of the action boundary of each new energy power generation equipment in a voltage abnormal fluctuation scenario.

5. The adaptive distributed renewable energy power generation equipment collaborative control method according to claim 4, characterized in that: The step of detecting the initial action space range comprises: The initial action space range is segmented using a preset time window length, and the time series variation characteristics of the power regulation range and the temperature limit condition fluctuation data within each time window are extracted; Analyzing the dynamic coupling relationship between the time series variation characteristics of the power adjustment range and the temperature limit condition fluctuation data through a linear regression algorithm, and calculating the dynamic variation trend coefficients of the two; The dynamic change trend coefficient is compared with a preset narrow range threshold, and if the dynamic change trend coefficient is less than the narrow range threshold, it is determined that the initial action space range is in a narrow range.

6. The adaptive distributed renewable energy power generation equipment collaborative control method according to claim 1, characterized in that: The step of obtaining the equipment operation reward parameter of the new energy power generation equipment by using a weighted method based on the data sampling frequency and the proportion of abnormal scenarios in the experience extended data set includes: By empirically expanding the data sampling frequency and the proportion of abnormal scenarios in the dataset, the initial reward function value is calculated using the weighted sum method; According to the mapping relationship between the distribution density of the initial reward function value and the proportion of abnormal scenes, a density peak clustering algorithm is used to obtain the reward interval; In each reward interval, the penalty factor benchmark value is calculated according to the load deviation between the equipment load factor and the preset load threshold, and the penalty factor benchmark value is smoothed by the sliding average method to obtain a smoothed dynamic cost penalty distribution; Based on the disturbance simulation amplitude, the dynamic cost penalty distribution in the high disturbance reward interval is compensated and corrected to obtain the corrected penalty distribution; The modified penalty distribution is inversely weighted superimposed on the initial reward function value, and normalized to obtain the equipment operation reward parameter of the new energy power generation equipment.

7. The adaptive distributed renewable energy power generation equipment collaborative control method according to claim 1, characterized in that: The step of extracting state-response data pairs having the same time decay coefficient from a pre-built distributed shared experience pool based on the device operation reward parameter to form a shared experience correction set comprises: Extracting a reference time decay coefficient from the device operation reward parameter as a matching reference, and matching the reference time decay coefficient with the time decay coefficient of each data pair in the distributed shared experience pool to extract a state-response data candidate set having the same time decay coefficient; Calculate the response delay difference between the action execution response time of each state-response data pair in the state-response data candidate set and the actual execution timestamp of the device response; Based on the response delay difference, a sliding window average method is used to align the timestamps in the state-response data candidate set to obtain a time-synchronized state-response data set; The state-response data set is divided into different time windows, and the state-response data of multiple renewable energy power generation equipment in the same time window are fused to construct a time-series continuous shared experience correction set.

8. The adaptive distributed renewable energy power generation equipment collaborative control method according to claim 1, characterized in that: The step of evaluating the probability of conflict in action selection of each new energy power generation device based on the shared experience correction set and the state feature representation comprises: Calculating the abnormal cumulative value of each new energy power generation device according to the state-response data in the shared experience correction set, and identifying high conflict risk intelligent agents from the new energy power generation devices according to the abnormal cumulative value; The voltage abnormal fluctuation amplitude is used as the main sorting basis, and the time decay coefficient is used as the secondary sorting basis. The entropy weight method is used to calculate the action priority of each high conflict risk intelligent agent to obtain the action priority intelligent agent sequence; Based on the action priority agent sequence, the overlapping ratio of power adjustment ranges of high-conflict risk agents with adjacent action priorities is calculated, and the potential conflict boundary is determined according to the overlapping ratio of power adjustment ranges; Map the real-time action instructions to the power adjustment interval of the potential conflict boundary, and count the frequency and magnitude of action instructions crossing the boundary for each high conflict risk agent; The conflict intensity index is calculated according to the product of the action instruction out-of-bounds frequency and the action instruction out-of-bounds amplitude, and the conflict intensity index is normalized to obtain the action selection conflict probability of each new energy power generation equipment.

9. The adaptive distributed renewable energy power generation equipment collaborative control method according to claim 1, characterized in that: The step of selecting the conflict probability according to the action and iteratively determining the multi-device collaborative control optimization strategy after conflict suppression using a gradient optimization algorithm comprises: According to the conflict probability of action selection of each renewable energy power generation equipment and the real-time state characteristics of the power grid, the corresponding initial disturbance frequency value is determined; Taking minimizing the conflict probability and maximizing the power output efficiency as the optimization goals, based on the optimization goals, using the gradient optimization algorithm to calculate the partial derivatives of the initial disturbance frequency value and the power adjustment range, and obtaining the gradient optimization direction; The disturbance frequency value and power adjustment range are gradually adjusted according to the gradient optimization direction. Through iterative calculation until the optimization target converges, the multi-device collaborative control optimization strategy after conflict suppression is output.

10. An adaptive distributed new energy power generation equipment collaborative control system, characterized in that: The system comprises: A data acquisition module is used to extract the status characteristic representation of each new energy power generation equipment from the real-time collected operating parameters of the new energy power generation equipment; A boundary estimation module, used to determine the estimated action boundary value of each new energy power generation equipment in the voltage abnormal fluctuation scenario by using a deep reinforcement learning algorithm based on the state feature representation; A data expansion module, used for extracting device state-action data pairs from a historical operation database according to the action boundary estimation value, and generating an empirical expansion data set by simulating different disturbance amplitudes according to the device state-action data pairs; The reward analysis module is used to obtain the equipment operation reward parameters of the new energy power generation equipment using a weighted method based on the data sampling frequency and the proportion of abnormal scenarios in the experience-expanded data set; An experience correction module, for extracting state-response data pairs with the same time decay coefficient from a pre-built distributed shared experience pool based on the device operation reward parameter to form a shared experience correction set; The strategy generation module is used to evaluate the action selection conflict probability of each new energy power generation equipment based on the shared experience correction set and the state feature representation, and according to the action selection conflict probability, use the gradient optimization algorithm to iteratively determine the multi-device collaborative control optimization strategy after conflict suppression.

Citation Information

Patent Citations

  • Continuous reinforcement learning incomplete information game method and device based on afterward review and progressive extension

    CN114048834A

  • Underwater detector cluster adaptive detection method and system based on distributed reinforcement learning

    CN119204155A

  • A dynamic optimization method for industrial automation system based on 5G private network

    CN119743772A

  • Conflict resolution mechanism for managing calendar events with a mobile communication device

    US20080114716A1

  • Deep reinforcement learning-based random access method for low earth orbit satellite network and terminal for the operation

    US20230189353A1

Cited By

  • Dynamic evaluation method and system for electric energy quality of new energy station

    CN120703497A

  • Distributed new energy cooperative control method and system

    CN121216625A

  • Photovoltaic panel angle self-adaptive adjusting system

    CN121560083A

  • Two-dimensional code equipment adaptive control method, system and equipment based on Internet of Things, and medium

    CN121995740A

  • Distributed energy agent regulation and control method and system based on deep reinforcement learning

    CN122001026A