Microgrid Reactive Power Optimization Method and System Based on Deep Learning

By adopting deep learning-based methods in microgrid reactive power optimization, and using comparative learning of active and negative reactive power configuration supervision data, the problems of complex calculations and difficult to adapt to dynamic changes in the prior art are solved, and more efficient and accurate reactive power optimization is achieved.

CN119787386BActive Publication Date: 2025-06-20KUNMING AUTOMATION WHOLE SET OF EQUIP BUSINESS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510266962.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-20
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

The prior art is complex in microgrid reactive optimization, is sensitive to initial parameters, and is difficult to adapt to dynamic changes of microgrids.

Method used

Using a deep learning-based method, model parameters are learned and optimized by obtaining candidate microgrid reactive power optimization models and training data sequences containing multiple sample learning data, and using the comparative learning mechanism between active reactive power configuration supervision data and passive reactive power configuration supervision data.

Benefits of technology

It significantly improves the performance in the reactive power optimization task of microgrid, avoids model overfitting, enhances the generalization ability of the model, and improves the accuracy and efficiency of reactive power optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119787386B_ABST
    Figure CN119787386B_ABST
Patent Text Reader

Abstract

The present invention provides a reactive power optimization method and system for a microgrid based on deep learning. By introducing a contrastive learning mechanism for positive reactive power configuration supervision data and negative reactive power configuration supervision data, the direct training benefits of each example learning data are considered, and an evaluation of the cyclic training reinforcement benefits is also introduced. By comparing the improvement in the prediction ability of the model among different training groups, overfitting of the model is effectively avoided, and the generalization ability of the model is promoted. In particular, by setting different fusion coefficients for the first cyclic goodness-of-fit value and the second cyclic goodness-of-fit value, the present invention further optimizes the learning weights of the model for positive and negative reactive power configurations, so that the generated target microgrid reactive power optimization model can more accurately reflect the actual requirements of microgrid operation, improving the accuracy and efficiency of reactive power optimization. Finally, the method can quickly generate efficient reactive power optimization configuration decisions for any given target microgrid operation data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and in particular, to a reactive power optimization method and system for a microgrid based on deep learning. Background Art

[0002] With the transformation of the energy structure and the rapid development of distributed energy technologies, the microgrid, as a small-scale power system integrating renewable energy, energy storage systems, and loads, has shown great potential in improving energy utilization efficiency, enhancing grid flexibility, and reliability. However, the reactive power optimization problem of the microgrid has always been one of the key factors restricting its efficient operation. Traditional reactive power optimization methods mostly rely on mathematical optimization algorithms, such as particle swarm optimization, genetic algorithms, etc. Although these methods can achieve the optimal configuration of reactive power compensation devices to a certain extent, they often have limitations such as high computational complexity, sensitivity to initial parameters, and difficulty in adapting to the dynamic changes of the microgrid. Summary of the Invention

[0003] In view of the problems mentioned above, in combination with the first aspect of the present invention, embodiments of the present invention provide a reactive power optimization method for a microgrid based on deep learning, and the method includes:

[0004] Obtain a candidate reactive power optimization model for the microgrid and a training data sequence including a plurality of sample learning data; the sample learning data includes sample microgrid operation data, negative reactive power configuration supervision data of the sample microgrid operation data, and positive reactive power configuration supervision data of the sample microgrid operation data;

[0005] For each sample learning data in the training data sequence, when performing model parameter learning based on the sample learning data in the target training stage, determine the target training reinforcement benefit of the positive reactive power configuration supervision data in the sample learning data compared to the negative reactive power configuration supervision data;

[0006] Determine a first loop goodness-of-fit value for the updated microgrid reactive power optimization model corresponding to the sample learning data to generate positive reactive power configuration supervision data, and a second loop goodness-of-fit value for the updated microgrid reactive power optimization model corresponding to the sample learning data to generate negative reactive power configuration supervision data; the updated microgrid reactive power optimization model corresponding to the sample learning data is generated by performing model parameter learning based on the previous sample training group of the sample training group corresponding to the sample learning data; in the process of the first model parameter learning, the deep learning target of the first sample training group is the candidate microgrid reactive power optimization model;

[0007] Perform a fusion calculation on the first-cycle goodness-of-fit value and the second-cycle goodness-of-fit value to determine the cycle training reinforcement benefit of the sample learning data. When it is determined that the training error in the target training phase no longer decreases based on the cycle training reinforcement benefit and the target training reinforcement benefit corresponding to each sample learning data respectively, terminate the model parameter learning and generate the target microgrid reactive power optimization model corresponding to the candidate microgrid reactive power optimization model; the fusion coefficient of the first-cycle goodness-of-fit value is less than the fusion coefficient of the second-cycle goodness-of-fit value;

[0008] Obtain the target microgrid operation data to be analyzed, and load the target microgrid operation data into the target microgrid reactive power optimization model to generate reactive power optimization configuration decision data for the target microgrid operation data.

[0009] In a possible implementation manner of the first aspect, the determining the first-cycle goodness-of-fit value of the updated microgrid reactive power optimization model corresponding to the sample learning data to generate positive reactive power configuration supervision data includes:

[0010] Obtain multiple candidate positive reactive power configuration supervision data of the sample learning data;

[0011] For each candidate positive reactive power configuration supervision data, determine the first-cycle goodness-of-fit value of the updated microgrid reactive power optimization model corresponding to the sample learning data to generate the candidate positive reactive power configuration supervision data;

[0012] Perform a mean calculation on the first-cycle goodness-of-fit values corresponding to each candidate positive reactive power configuration supervision data to generate the first-cycle goodness-of-fit value of the updated microgrid reactive power optimization model corresponding to the sample learning data to generate positive reactive power configuration supervision data.

[0013] In a possible implementation manner of the first aspect, the method further includes:

[0014] Determine the positive supervision training reinforcement benefit of the updated microgrid reactive power optimization model corresponding to the sample learning data compared with the candidate microgrid reactive power optimization model, and the negative supervision training reinforcement benefit of the updated microgrid reactive power optimization model corresponding to the sample learning data compared with the candidate microgrid reactive power optimization model;

[0015] Determine the model training comparison loss corresponding to the sample learning data based on the positive supervision training reinforcement benefit and the negative supervision training reinforcement benefit;

[0016] When it is determined that the training error in the target training stage no longer continues to decrease based on the cyclic training reinforcement benefit and the target training reinforcement benefit corresponding to each of the sample learning data respectively, terminating the learning of the model parameters and generating the target microgrid reactive power optimization model corresponding to the candidate microgrid reactive power optimization model, includes:

[0017] When it is determined that the training error in the target training stage no longer continues to decrease based on the cyclic training reinforcement benefit, the target training reinforcement benefit, and the model training comparison loss corresponding to each of the sample learning data respectively, terminating the learning of the model parameters and generating the target microgrid reactive power optimization model corresponding to the candidate microgrid reactive power optimization model.

[0018] In a possible implementation manner of the first aspect, determining the negative supervised training reinforcement benefit of the updated microgrid reactive power optimization model corresponding to the sample learning data compared with the candidate microgrid reactive power optimization model, includes:

[0019] Determining the negative reactive power configuration supervised data training reinforcement benefits of the updated microgrid reactive power optimization model compared with the candidate microgrid reactive power optimization model respectively under each of the sample learning data in the sample training group corresponding to the sample learning data;

[0020] Calculating the mean value of each of the negative reactive power configuration supervised data training reinforcement benefits to determine the negative supervised training reinforcement benefit of the updated microgrid reactive power optimization model corresponding to the sample learning data compared with the candidate microgrid reactive power optimization model.

[0021] In a possible implementation manner of the first aspect, the method further includes:

[0022] For each of the sample learning data, determining the reference training error corresponding to the sample learning data according to the cyclic training reinforcement benefit and the target training reinforcement benefit of the sample learning data;

[0023] Determining the first fusion coefficient of the reference training error under the sample learning data and the second fusion coefficient of the model training comparison loss under the sample learning data;

[0024] According to the first fusion coefficient and the second fusion coefficient, performing a fusion calculation on the reference training error and the model training comparison loss of the sample learning data to determine the training error corresponding to the sample learning data;

[0025] When the calculation results of the training errors corresponding to each of the sample learning data no longer continue to decrease, determining that the training error in the target training stage no longer continues to decrease.

[0026] In a possible implementation of the first aspect, determining the first fusion coefficient of the reference training error under the sample learning data includes:

[0027] Determining the updated training reinforcement benefit of the positive reactive power configuration supervision data compared to the negative reactive power configuration supervision data after completing model parameter learning under the sample learning data;

[0028] Based on the first difference between the updated training reinforcement benefit and the target training reinforcement benefit, determining the first fusion coefficient of the reference training error under the sample learning data; the first fusion coefficient has a negative correlation with the first difference;

[0029] Wherein, based on the first difference between the updated training reinforcement benefit and the target training reinforcement benefit, determining the first fusion coefficient of the reference training error under the sample learning data includes:

[0030] Based on the updated training reinforcement benefit and the target training reinforcement benefit corresponding to each sample learning data in the sample training group corresponding to the sample learning data, determining the first difference corresponding to each sample learning data;

[0031] Calculating the mean value of the first differences corresponding to each sample learning data to determine the first threshold value matching the sample learning data;

[0032] When the first difference corresponding to the sample learning data is not greater than the first threshold value, determining the first fusion coefficient of the reference training error under the sample learning data according to the negative correlation relationship.

[0033] In a possible implementation of the first aspect, determining the second fusion coefficient of the model training comparison loss under the sample learning data includes:

[0034] Determining the training reinforcement benefit of the negative reactive power configuration supervision data of the updated microgrid reactive power optimization model corresponding to the sample learning data compared to the candidate microgrid reactive power optimization model;

[0035] Based on the second difference between the positive supervision training reinforcement benefit and the training reinforcement benefit of the negative reactive power configuration supervision data, determining the second fusion coefficient of the model training comparison loss under the sample learning data; the second fusion coefficient has a negative correlation with the second difference;

[0036] Wherein, based on the second difference between the positive supervision training reinforcement benefit and the training reinforcement benefit of the negative reactive power configuration supervision data, determining the second fusion coefficient of the model training comparison loss under the sample learning data includes:

[0037] Determine the second difference corresponding to each of the sample learning data according to the positive supervised training reinforcement benefit and the negative reactive power configuration supervised data training reinforcement benefit corresponding to each of the sample learning data in the sample training group corresponding to the sample learning data.

[0038] Perform a mean calculation on the second differences corresponding to each of the sample learning data to determine a second threshold value matching the sample learning data.

[0039] When the second difference corresponding to the sample learning data is not greater than the second threshold value, determine a second fusion coefficient of the model training comparison loss under the sample learning data according to the negative correlation relationship.

[0040] In a possible implementation manner of the first aspect, the process of obtaining a training data sequence including a plurality of sample learning data includes:

[0041] Obtain the reactive power optimization requirements of the microgrid and a plurality of sample microgrid operation data;

[0042] For each of the sample microgrid operation data, determine the demand index data associated with the reactive power optimization requirements of the microgrid; the demand index data includes the sample microgrid operation data;

[0043] Load the demand index data into a pre-trained deep learning network, and use the generation of the deep learning network as the negative reactive power configuration supervised data of the sample microgrid operation data;

[0044] Generate the positive reactive power configuration supervised data of the sample microgrid operation data according to the expert optimization instruction generated for the negative reactive power configuration supervised data of the sample microgrid operation data;

[0045] Configure sample learning data including the sample microgrid operation data, the positive reactive power configuration supervised data of the sample microgrid operation data, and the negative reactive power configuration supervised data of the sample microgrid operation data, and generate a training data sequence including a plurality of sample learning data.

[0046] In a possible implementation manner of the first aspect, the process of obtaining a candidate microgrid reactive power optimization model includes:

[0047] Obtain an initialized neural network model;

[0048] According to the training data sequence, configure a basic training data sequence including a plurality of basic sample learning data; each of the basic sample learning data includes sample microgrid operation data and the positive reactive power configuration supervised data or the negative reactive power configuration supervised data of the sample microgrid operation data;

[0049] Iteratively update the neural network model according to the basic training data sequence to generate a candidate microgrid reactive power optimization model.

[0050] On the other hand, an embodiment of the present invention further provides a microgrid reactive power optimization system based on deep learning, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.

[0051] Based on the above aspects, the embodiment of the present application significantly improves the performance of the model in the microgrid reactive power optimization task by introducing a comparative learning mechanism for positive reactive power configuration supervision data and negative reactive power configuration supervision data. During the training process, this method not only considers the direct training benefits of each example learning data, but also introduces an evaluation of the cyclic training reinforcement benefits. By comparing the improvement of the model's prediction ability among different training groups, overfitting of the model is effectively avoided, and the generalization ability of the model is promoted. In particular, by setting different fusion coefficients for the first cyclic goodness-of-fit value and the second cyclic goodness-of-fit value, the present invention further optimizes the learning weights of the model for positive and negative reactive power configurations, so that the generated target microgrid reactive power optimization model can more accurately reflect the actual needs of microgrid operation, improving the accuracy and efficiency of reactive power optimization. Finally, this method can quickly generate efficient reactive power optimization configuration decisions for any given target microgrid operation data, thus contributing to the stable operation and energy efficiency improvement of the microgrid. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is a schematic flowchart of the execution process of the microgrid reactive power optimization method based on deep learning provided by an embodiment of the present invention.

[0053] Figure 2 is a schematic diagram of the hardware architecture of the microgrid reactive power optimization system based on deep learning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The present invention will be specifically described below with reference to the accompanying drawings of the specification. Figure 1 is a schematic flowchart of the microgrid reactive power optimization method based on deep learning provided by an embodiment of the present invention. The microgrid reactive power optimization method based on deep learning will be introduced in detail below.

[0055] Step S110: Obtain a candidate microgrid reactive power optimization model and a training data sequence containing multiple sample learning data. The sample learning data includes sample microgrid operation data, negative reactive power configuration supervision data of the sample microgrid operation data, and positive reactive power configuration supervision data of the sample microgrid operation data.

[0056] In this embodiment, in a microgrid system, it is assumed that there is a large industrial park microgrid. This microgrid includes multiple distributed power sources (such as solar photovoltaic panels, small wind turbines), different types of loads (industrial electrical equipment, office area electrical equipment, etc.), and related power transmission and conversion equipment (transformers, switchgear, etc.).

[0057] For obtaining the candidate microgrid reactive power optimization model, select an initialized neural network model from an existing neural network model library. This neural network model may be a multi-layer perceptron (MLP) structure, which has an input layer, several hidden layers, and an output layer. The number of nodes in the input layer may be determined according to the number of characteristics of the microgrid operation data. For example, characteristics such as the voltage, current, power factor, active power demand, and reactive power demand of the load in the microgrid are used as inputs to the input layer.

[0058] Next, obtain the training data sequence. Multiple sets of sample microgrid operation data at different time periods are collected. For example, from 9 am to 10 am on a certain day, when the solar photovoltaic panels are generating electricity under medium light intensity, the industrial electrical equipment is in normal production state, and some of the office area electrical equipment is turned on. The corresponding microgrid operation data includes the real-time voltage of the entire microgrid being 380V, current being 50A, power factor being 0.85, the output power of each distributed power source, and the power demand of the load, etc.

[0059] For determining the negative reactive power configuration supervision data, load these sample microgrid operation data as demand index data into a pre-trained deep learning network. This deep learning network may be one that was preliminarily trained for the microgrid reactive power configuration problem before. Taking the sample from 9 am to 10 am as an example, the deep learning network, according to the input microgrid operation data and the rules and algorithms it has pre-learned, outputs a reactive power configuration plan. For example, the reactive power compensation amount set for a certain reactive power compensation device is Q1. This reactive power configuration plan is used as the negative reactive power configuration supervision data.

[0060] Then, positive reactive power configuration supervision data is generated. According to the expert optimization instructions generated for the above-mentioned negative reactive power configuration supervision data, experts optimize the negative reactive power configuration supervision data based on their own experience and a deeper understanding of the microgrid system, considering objectives such as reducing grid losses and improving power quality. For example, experts believe that under the current operating conditions, by adjusting the parameters of a certain reactive power compensation device, the reactive power compensation amount can be adjusted to Q2 (Q2 is different from Q1 and more in line with the optimization objective), and the reactive power configuration scheme corresponding to this Q2 is used as the positive reactive power configuration supervision data.

[0061] In this way, for each sample microgrid operation data, corresponding negative reactive power configuration supervision data and positive reactive power configuration supervision data are generated, thus configuring into sample learning data. Multiple such sample learning data form a training data sequence.

[0062] Step S120: For each sample learning data in the training data sequence, when performing model parameter learning based on the sample learning data in the target training stage, determine the target training reinforcement benefit of the positive reactive power configuration supervision data compared to the negative reactive power configuration supervision data in the sample learning data.

[0063] Continuing with the above-mentioned industrial park microgrid as an example. In the target training stage, each sample learning data in the training data sequence is used for model parameter learning. For a certain sample learning data, such as the sample from 9 am to 10 am mentioned before.

[0064] First, input the sample microgrid operation data into the model currently being trained (which may be a candidate microgrid reactive power optimization model or an updated model after a certain number of iterations at the initial stage of training). The model predicts the results of positive reactive power configuration and negative reactive power configuration based on the input operation data. Assume that the reactive power compensation amount predicted by the model according to the input features corresponding to the negative reactive power configuration supervision data is Q1', and the reactive power compensation amount predicted according to the input features corresponding to the positive reactive power configuration supervision data is Q2'.

[0065] The determination of the target training reinforcement benefit can be considered from multiple aspects. For example, it can be measured from aspects such as the accuracy of reactive power compensation and the contribution to the improvement of the overall performance of the microgrid. From the perspective of the accuracy of reactive power compensation, a measurement index can be defined, such as the error rate. Suppose the error rate of the supervision data according to the negative reactive power configuration is E1 = |Q1 - Q1'| / Q1, and the error rate of the supervision data according to the positive reactive power configuration is E2 = |Q2 - Q2'| / Q2. Then the target training reinforcement benefit of the positive reactive power configuration supervision data compared with the negative reactive power configuration supervision data in terms of the accuracy of reactive power compensation can be expressed as ΔE = E1 - E2. If ΔE is greater than 0, it means that the positive reactive power configuration supervision data has a better effect than the negative reactive power configuration supervision data in terms of the accuracy of reactive power compensation under this sample learning data, and this ΔE is an embodiment of the target training reinforcement benefit.

[0066] From the perspective of the improvement of the overall performance of the microgrid, the impacts of reactive power compensation on aspects such as voltage stability and grid loss can be considered. Suppose under the negative reactive power configuration, the voltage fluctuation range of the microgrid is ΔV1, and under the positive reactive power configuration, the voltage fluctuation range is ΔV2. If ΔV1 is greater than ΔV2, it indicates that the positive reactive power configuration has a better improvement effect on voltage stability. The benefits in aspects such as the accuracy of reactive power compensation and the improvement of the overall performance of the microgrid can be combined according to a certain weight to obtain the target training reinforcement benefit of the positive reactive power configuration supervision data compared with the negative reactive power configuration supervision data under this sample learning data.

[0067] Step S130, determine the first cycle goodness-of-fit value of the updated microgrid reactive power optimization model corresponding to the sample learning data to generate positive reactive power configuration supervision data, and the second cycle goodness-of-fit value of the updated microgrid reactive power optimization model corresponding to the sample learning data to generate negative reactive power configuration supervision data. The updated microgrid reactive power optimization model corresponding to the sample learning data is generated by learning model parameters based on the previous sample training group of the sample training group corresponding to the sample learning data. In the process of the first model parameter learning, the deep learning target of the first sample training group is the candidate microgrid reactive power optimization model.

[0068] Still taking the industrial park microgrid as an example. Suppose the training data sequence is divided into multiple sample training groups. For the sample training group where a certain sample learning data is located, such as the nth sample training group.

[0069] First, determine the first-cycle goodness-of-fit value. Obtain multiple candidate positive reactive power configuration supervision data for this sample learning data. These candidate positive reactive power configuration supervision data may be generated by making some minor perturbations to the positive reactive power configuration supervision data or based on different algorithms. For example, for the positive reactive power configuration supervision data Q2 in the sample from 9 am to 10 am mentioned above, by fine-tuning the parameters of the reactive power compensation device, candidate positive reactive power configuration supervision data such as Q21, Q22, Q23, etc. can be obtained.

[0070] For each candidate positive reactive power configuration supervision data, input the sample microgrid operation data into the updated microgrid reactive power optimization model generated by learning the model parameters according to the (n - 1)-th sample training group (in the first sample training group, the deep learning target is the candidate microgrid reactive power optimization model). Assume that for the candidate positive reactive power configuration supervision data Q21, the predicted reactive power compensation amount output by the model is Q21'. Then, the first-cycle goodness-of-fit value of the updated microgrid reactive power optimization model corresponding to this sample learning data for generating Q21 can be calculated based on the difference between the predicted value and the actual value. For example, the mean square error (MSE) formula can be used, MSE = (Q21 - Q21')². Similarly, for candidate positive reactive power configuration supervision data such as Q22 and Q23, the corresponding first-cycle goodness-of-fit values are also calculated. Finally, the mean of these first-cycle goodness-of-fit values is calculated to obtain the first-cycle goodness-of-fit value of the updated microgrid reactive power optimization model corresponding to this sample learning data for generating positive reactive power configuration supervision data.

[0071] For the determination of the second-cycle goodness-of-fit value, the process is similar. Similar operations are also performed on the negative reactive power configuration supervision data. Assume that the negative reactive power configuration supervision data is Q1, and multiple candidate negative reactive power configuration supervision data Q11, Q12, Q13, etc. can be obtained (generated by similar perturbations or different algorithms). Input the sample microgrid operation data into the updated microgrid reactive power optimization model, calculate the goodness-of-fit values corresponding to each candidate negative reactive power configuration supervision data (also using methods such as MSE), and finally calculate the mean to obtain the second-cycle goodness-of-fit value of the updated microgrid reactive power optimization model corresponding to this sample learning data for generating negative reactive power configuration supervision data.

[0072] Step S140: Perform a fusion calculation on the first cycle goodness-of-fit value and the second cycle goodness-of-fit value to determine the cycle training reinforcement benefit of the sample learning data. When it is determined that the training error in the target training stage no longer continues to decrease based on the cycle training reinforcement benefit and the target training reinforcement benefit corresponding to each sample learning data respectively, terminate the model parameter learning and generate the target microgrid reactive power optimization model corresponding to the candidate microgrid reactive power optimization model. The fusion coefficient of the first cycle goodness-of-fit value is less than the fusion coefficient of the second cycle goodness-of-fit value.

[0073] In the scenario of the industrial park microgrid, continue with the previous example. Assume that for a certain sample learning data, the first cycle goodness-of-fit value F1 and the second cycle goodness-of-fit value F2 have been calculated. Since the fusion coefficient of the first cycle goodness-of-fit value is less than the fusion coefficient of the second cycle goodness-of-fit value, assume that the fusion coefficient of the first cycle goodness-of-fit value is α and the fusion coefficient of the second cycle goodness-of-fit value is β (α < β and α + β = 1).

[0074] Then the cycle training reinforcement benefit of this sample learning data can be calculated by the formula: cycle training reinforcement benefit = α * F1 + β * F2.

[0075] During the entire training process, such calculations are performed on each sample learning data in the training data sequence. Then, the training error is determined based on the cycle training reinforcement benefit and the target training reinforcement benefit corresponding to each sample learning data respectively.

[0076] For example, for each sample learning data, the calculation method of the training error can be defined as: training error = |cycle training reinforcement benefit - target training reinforcement benefit|. As the model parameter learning progresses, the training error corresponding to each sample learning data is continuously calculated. When it is found that the calculation results of the training errors corresponding to all sample learning data no longer continue to decrease, it indicates that the model has converged to a better state. At this time, terminate the model parameter learning. Finally, use the model at this time as the target microgrid reactive power optimization model corresponding to the candidate microgrid reactive power optimization model. This target microgrid reactive power optimization model has high accuracy and reliability in making decisions on microgrid reactive power optimization configuration.

[0077] Step S150: Obtain the target microgrid operation data to be analyzed, and load the target microgrid operation data into the target microgrid reactive power optimization model to generate the reactive power optimization configuration decision data of the target microgrid operation data.

[0078] Suppose in the industrial park microgrid, the operation data of the target microgrid to be analyzed from 3 pm to 4 pm on a new day is obtained. At this time, the light intensity of the solar photovoltaic panels may change due to cloud cover, and the load of industrial electrical equipment also changes. For example, the voltage becomes 375V, the current is 45A, the power factor is 0.83, and the output power of each distributed power source and the power demand of the load and other data have corresponding changes.

[0079] Load this target microgrid operation data into the previously obtained target microgrid reactive power optimization model. The target microgrid reactive power optimization model analyzes and processes the input target microgrid operation data according to the parameters and algorithms it has learned internally. The model will comprehensively consider various factors such as voltage stability, grid loss, and the characteristics of reactive power compensation devices based on the current operating state of the microgrid, and output the reactive power optimization configuration decision data for this target microgrid operation data. For example, the model may output that the reactive power compensation amount of a certain reactive power compensation device should be adjusted to Q3, and the control parameters of other related devices are also set accordingly, so as to achieve the reactive power optimization configuration of the microgrid during this time period and improve the operating efficiency and power quality of the microgrid.

[0080] Based on the above steps, the embodiment of the present application significantly improves the performance of the model in the microgrid reactive power optimization task by introducing a comparative learning mechanism of positive reactive power configuration supervision data and negative reactive power configuration supervision data. During the training process, this method not only considers the direct training benefits of each sample learning data, but also introduces the evaluation of cyclic training reinforcement benefits. By comparing the improvement of the model's prediction ability among different training groups, it effectively avoids model overfitting and promotes the enhancement of the model's generalization ability. In particular, by setting different fusion coefficients for the first cyclic goodness-of-fit value and the second cyclic goodness-of-fit value, the present invention further optimizes the learning weights of the model for positive and negative reactive power configurations, enabling the generated target microgrid reactive power optimization model to more accurately reflect the actual needs of microgrid operation, improving the accuracy and efficiency of reactive power optimization. Finally, this method can quickly generate efficient reactive power optimization configuration decisions for any given target microgrid operation data, thus contributing to the stable operation and energy efficiency improvement of the microgrid.

[0081] In a possible implementation manner, step S130 includes:

[0082] Step S131, obtaining multiple candidate positive reactive power configuration supervision data of the sample learning data.

[0083] Step S132, for each of the candidate positive reactive power configuration supervision data, determining the first cyclic goodness-of-fit value of the updated microgrid reactive power optimization model corresponding to the sample learning data for generating the candidate positive reactive power configuration supervision data.

[0084] In step S133, calculate the mean value of the first cycle goodness-of-fit values corresponding to each of the candidate positive reactive power configuration supervision data, and generate the first cycle goodness-of-fit value of the updated microgrid reactive power optimization model corresponding to the sample learning data for generating the positive reactive power configuration supervision data.

[0085] In this embodiment, in the scenario of an industrial park microgrid, consider the situation from 9:00 to 10:00 in the morning in the sample learning data. To determine the first cycle goodness-of-fit value of the updated microgrid reactive power optimization model corresponding to the sample learning data for generating the positive reactive power configuration supervision data, the first operation is to obtain multiple candidate positive reactive power configuration supervision data of the sample learning data. In this specific microgrid operation period, the original positive reactive power configuration supervision data is the optimization configuration data for the reactive power compensation device generated based on the expert optimization instructions. To obtain multiple candidate positive reactive power configuration supervision data, various methods can be used. For example, adjust the original positive reactive power configuration supervision data according to different reactive power compensation strategy adjustment algorithms. Assume that the compensation amount of the reactive power compensation device in the original positive reactive power configuration supervision data is set to Q2, adjust it to Q21 through one adjustment algorithm, and then adjust it to Q22 through another algorithm, and so on to obtain multiple different candidate positive reactive power configuration supervision data.

[0086] For each candidate positive reactive power configuration supervision data, determine the first cycle goodness-of-fit value of the updated microgrid reactive power optimization model corresponding to the sample learning data for generating this candidate positive reactive power configuration supervision data. Input the sample microgrid operation data from 9:00 to 10:00 in the morning (including data such as voltage, current, power factor, distributed power source output power, and load power demand) into the updated microgrid reactive power optimization model generated by learning the model parameters in the previous sample training group corresponding to the sample training group corresponding to the sample learning data. Taking the candidate positive reactive power configuration supervision data Q21 as an example, when the sample microgrid operation data is input into the updated microgrid reactive power optimization model, the model will output a predicted reactive power compensation amount Q21' according to its internal parameters and algorithm structure. At this time, the first cycle goodness-of-fit value can be determined according to the difference between the predicted value Q21' and the actual candidate positive reactive power configuration supervision data Q21. A common determination method is to use the mean square error (MSE) calculation method, that is, calculate the value of (Q21 - Q21')² as the first cycle goodness-of-fit value corresponding to Q21. In the same way, for the candidate positive reactive power configuration supervision data Q22, etc., also input the sample microgrid operation data into the updated microgrid reactive power optimization model, obtain the corresponding predicted values, and then calculate the mean square error with the actual candidate positive reactive power configuration supervision data to determine their respective corresponding first cycle goodness-of-fit values.

[0087] Finally, calculate the mean of the first-cycle goodness-of-fit values corresponding to each candidate positive reactive power configuration supervision data, and generate the first-cycle goodness-of-fit value of the updated microgrid reactive power optimization model corresponding to the sample learning data for generating positive reactive power configuration supervision data. Suppose the first-cycle goodness-of-fit value calculated for Q21 is F21, the first-cycle goodness-of-fit value calculated for Q22 is F22, and so on. Sum all these first-cycle goodness-of-fit values F21, F22, etc., and then divide by the number of candidate positive reactive power configuration supervision data. The result obtained is the first-cycle goodness-of-fit value of the updated microgrid reactive power optimization model corresponding to the sample learning data for generating positive reactive power configuration supervision data. This first-cycle goodness-of-fit value can reflect the fitting effect of the updated microgrid reactive power optimization model on positive reactive power configuration supervision data, and is of great significance in the subsequent model evaluation and optimization process. It helps to accurately measure the performance of the model under different reactive power configurations, thereby providing a more reliable decision-making basis for microgrid reactive power optimization.

[0088] In a possible implementation manner, the method further includes:

[0089] Step A110, determining the positive supervision training reinforcement benefit of the updated microgrid reactive power optimization model corresponding to the sample learning data compared to the candidate microgrid reactive power optimization model, and the negative supervision training reinforcement benefit of the updated microgrid reactive power optimization model corresponding to the sample learning data compared to the candidate microgrid reactive power optimization model.

[0090] Step A120, determining the model training comparison loss corresponding to the sample learning data based on the positive supervision training reinforcement benefit and the negative supervision training reinforcement benefit.

[0091] Step S140 includes: when it is determined that the training error in the target training stage no longer continues to decrease based on the cycle training reinforcement benefit, target training reinforcement benefit, and model training comparison loss corresponding to each sample learning data respectively, terminating the model parameter learning, and generating the target microgrid reactive power optimization model corresponding to the candidate microgrid reactive power optimization model.

[0092] In a possible implementation manner, step A110 includes:

[0093] Step A111, determining the negative reactive power configuration supervision data training reinforcement benefits corresponding to the updated microgrid reactive power optimization model compared to the candidate microgrid reactive power optimization model under each sample learning data in the sample training group corresponding to the sample learning data.

[0094] Step A112, calculate the mean value of the training reinforcement benefits of the negative reactive power configuration supervision data for each of the above, and determine the negative supervision training reinforcement benefit of the updated microgrid reactive power optimization model corresponding to the sample learning data compared to the candidate microgrid reactive power optimization model.

[0095] In this embodiment, first, consider the positive supervision training reinforcement benefit. For the situation of the sample learning data from 9:00 am to 10:00 am, during the operation of the microgrid in this period, the sample learning data includes sample microgrid operation data, negative reactive power configuration supervision data of the sample microgrid operation data, and positive reactive power configuration supervision data. When it comes to the positive supervision training reinforcement benefit of the updated microgrid reactive power optimization model compared to the candidate microgrid reactive power optimization model, it needs to be measured from multiple technical indicators. For example, consider aspects such as the improvement in reactive power compensation accuracy, the degree of improvement in voltage stability, and the effect of reducing grid losses.

[0096] In terms of reactive power compensation accuracy, input the sample microgrid operation data from 9:00 am to 10:00 am into the updated microgrid reactive power optimization model and the candidate microgrid reactive power optimization model respectively. Assume that under the positive reactive power configuration supervision data, the predicted value of the reactive power compensation amount output by the candidate microgrid reactive power optimization model is Q2', and the predicted value of the reactive power compensation amount output by the updated microgrid reactive power optimization model is Q2''. Measure the positive supervision training reinforcement benefit by calculating the error difference between the two and the reactive power compensation amount Q2 in the actual positive reactive power configuration supervision data. For example, the formula can be used: Positive supervision training reinforcement benefit (reactive power compensation accuracy) = |Q2 - Q2'| - |Q2 - Q2''|. If this value is greater than 0, it means that the updated microgrid reactive power optimization model has a reinforcement benefit compared to the candidate microgrid reactive power optimization model in terms of reactive power compensation accuracy.

[0097] For the degree of improvement in voltage stability, calculate the change in the voltage fluctuation range of the microgrid under the two models according to the reactive power configuration results output by the models. Assume that the voltage fluctuation range under the candidate microgrid reactive power optimization model is ΔV1, and the voltage fluctuation range under the updated microgrid reactive power optimization model is ΔV2. Positive supervision training reinforcement benefit (voltage stability) = ΔV1 - ΔV2. If this value is greater than 0, it indicates that the updated microgrid reactive power optimization model has a better performance in terms of voltage stability.

[0098] For the effect of reducing power grid losses, according to the relationship model between reactive power configuration and power grid losses, calculate the power grid loss values under the two models. Let the power grid loss under the candidate microgrid reactive power optimization model be L1, and the power grid loss under the updated microgrid reactive power optimization model be L2. The active supervision training reinforcement benefit (power grid loss) = L1 - L2. If it is greater than 0, it indicates that the updated microgrid reactive power optimization model has a positive reinforcement benefit in reducing power grid losses. Considering the benefits in these aspects and through a certain weight allocation (such as determining the weights according to the actual operation requirements of the microgrid), obtain the active supervision training reinforcement benefit of the updated microgrid reactive power optimization model compared with the candidate microgrid reactive power optimization model corresponding to the sample learning data.

[0099] Similarly, for the passive supervision training reinforcement benefit, determine the passive reactive power configuration supervision data training reinforcement benefits of the updated microgrid reactive power optimization model compared with the candidate microgrid reactive power optimization model under each sample learning data in the sample training group corresponding to the sample learning data. Taking multiple sample learning data in the sample training group as an example, for each sample learning data among them, such as the sample from 9 am to 10 am, input the sample microgrid operation data into the updated microgrid reactive power optimization model and the candidate microgrid reactive power optimization model. Assume that under the passive reactive power configuration supervision data, the predicted reactive power compensation value output by the candidate microgrid reactive power optimization model is Q1', and the predicted reactive power compensation value output by the updated microgrid reactive power optimization model is Q1''. Calculate the passive reactive power configuration supervision data training reinforcement benefit. For example, in terms of reactive power compensation accuracy, the passive reactive power configuration supervision data training reinforcement benefit (reactive power compensation accuracy) = |Q1 - Q1'| - |Q1 - Q1''|.

[0100] For voltage stability and power grid losses, calculate in a similar manner. For example, in terms of voltage stability, let the voltage fluctuation range under the candidate microgrid reactive power optimization model be ΔV3, and the voltage fluctuation range under the updated microgrid reactive power optimization model be ΔV4. The passive reactive power configuration supervision data training reinforcement benefit (voltage stability) = ΔV3 - ΔV4. In terms of power grid losses, let the power grid loss under the candidate microgrid reactive power optimization model be L3, and the power grid loss under the updated microgrid reactive power optimization model be L4. The passive reactive power configuration supervision data training reinforcement benefit (power grid loss) = L3 - L4. Then calculate the mean value of these passive reactive power configuration supervision data training reinforcement benefits under each sample learning data in the sample training group. Assume there are n sample learning data, and the calculated passive reactive power configuration supervision data training reinforcement benefits under each sample learning data are B1, B2,..., Bn respectively. Then the passive supervision training reinforcement benefit of the updated microgrid reactive power optimization model compared with the candidate microgrid reactive power optimization model corresponding to the sample learning data = (B1 + B2 +... + Bn) / n.

[0101] Based on the obtained positive supervision training reinforcement benefit and negative supervision training reinforcement benefit described above, determine the model training comparison loss corresponding to the example learning data. For example, the model training comparison loss can be defined in the form of a difference, that is, model training comparison loss = positive supervision training reinforcement benefit - negative supervision training reinforcement benefit. This model training comparison loss reflects the benefit difference between the positive and negative supervision training for updating the microgrid reactive power optimization model, which helps to more comprehensively evaluate the training effect of the model.

[0102] When it is determined that the training error in the target training stage no longer decreases based on the loop training reinforcement benefit, target training reinforcement benefit, and model training comparison loss corresponding to each example learning data respectively, terminate the model parameter learning and generate the target microgrid reactive power optimization model corresponding to the candidate microgrid reactive power optimization model. For each example learning data, the loop training reinforcement benefit is obtained through specific calculations before, and the target training reinforcement benefit is also determined based on the comparison of positive and negative reactive power configuration supervision data, plus the model training comparison loss. Taking a certain example learning data as an example, assuming the loop training reinforcement benefit is C, the target training reinforcement benefit is T, and the model training comparison loss is L, one way to determine the training error can be: training error = |C - T| + |L|. During the entire training process, continuously calculate this training error corresponding to each example learning data. As the model parameter learning progresses, when the training error corresponding to each example learning data no longer decreases, it indicates that the model has reached a relatively stable state, and at this time, terminate the model parameter learning. The finally obtained model is the target microgrid reactive power optimization model corresponding to the candidate microgrid reactive power optimization model. This target model will have high accuracy and reliability in the reactive power optimization configuration decision of the industrial park microgrid, can better adapt to the operation requirements of the microgrid, and improve the operation efficiency and stability of the microgrid.

[0103] In a possible implementation manner, the method further includes:

[0104] Step B110, for each of the example learning data, determine the reference training error corresponding to the example learning data according to the loop training reinforcement benefit and target training reinforcement benefit of the example learning data.

[0105] Step B120, determine the first fusion coefficient of the reference training error and the second fusion coefficient of the model training comparison loss under the example learning data.

[0106] Step B130: Based on the first fusion coefficient and the second fusion coefficient, perform a fusion calculation on the reference training error and the model training comparison loss of the sample learning data to determine the training error corresponding to the sample learning data.

[0107] Step B140: When the calculation results of the training errors corresponding to each of the sample learning data no longer continue to decrease, determine that the training error in the target training phase no longer continues to decrease.

[0108] In a possible implementation manner, step B120 includes:

[0109] Step B121: Determine the updated training reinforcement benefit of the positive reactive power configuration supervision data compared to the negative reactive power configuration supervision data after completing model parameter learning under the sample learning data.

[0110] Step B122: Based on the first difference between the updated training reinforcement benefit and the target training reinforcement benefit, determine the first fusion coefficient of the reference training error under the sample learning data. The first fusion coefficient has a negative correlation with the first difference.

[0111] Among them, step B122 includes:

[0112] Step B1221: Based on the updated training reinforcement benefits and the target training reinforcement benefits respectively corresponding to each sample learning data in the sample training group corresponding to the sample learning data, determine the first differences respectively corresponding to each sample learning data.

[0113] Step B122: Calculate the mean value of the first differences respectively corresponding to each sample learning data to determine the first threshold value matching the sample learning data.

[0114] Step B1223: When the first difference corresponding to the sample learning data is not greater than the first threshold value, determine the first fusion coefficient of the reference training error under the sample learning data according to the negative correlation.

[0115] In this embodiment, the operation of the microgrid from 9 am to 10 am in the example learning data is taken as an example. The cyclic training reinforcement benefit is obtained through a series of previous calculations, reflecting the performance of the model during the cyclic training process, while the target training reinforcement benefit is determined based on the performance of the positive reactive power configuration supervision data compared to the negative reactive power configuration supervision data during the target training stage. The determination method of the reference training error can be based on the difference relationship between the two. Assuming the cyclic training reinforcement benefit is CT and the target training reinforcement benefit is TT, then the reference training error can be calculated by the formula: reference training error = |CT - TT|. This reference training error can reflect the deviation degree of the example learning data during the training process from one perspective and is an important basis for subsequent calculations.

[0116] Next, determine the first fusion coefficient of the reference training error and the second fusion coefficient of the model training comparison loss under the example learning data. First, determine the updated training reinforcement benefit of the positive reactive power configuration supervision data compared to the negative reactive power configuration supervision data under the example learning data after completing the model parameter learning. For the example learning data from 9 am to 10 am, after completing the model parameter learning, input the example microgrid operation data into the updated model respectively, and obtain the corresponding output results for the positive reactive power configuration supervision data and the negative reactive power configuration supervision data. Assume the output result under the positive reactive power configuration supervision data is Q2', the output result under the negative reactive power configuration supervision data is Q1', the actual positive reactive power configuration supervision data is Q2, and the negative reactive power configuration supervision data is Q1. The updated training reinforcement benefit can be calculated by comprehensively considering aspects such as the accuracy of reactive power compensation, the impact on voltage stability, and grid losses. For example, in terms of the accuracy of reactive power compensation, the updated training reinforcement benefit (reactive power compensation accuracy) = |Q1 - Q1'| - |Q2 - Q2'|; in terms of voltage stability, calculate the corresponding benefit difference according to the impact of the reactive power configuration output by the model on the voltage fluctuation range; in terms of grid losses, calculate the benefit difference based on the relationship between the reactive power configuration and grid losses, and then comprehensively consider these aspects (through a certain weight distribution, and the weight is determined according to the microgrid operation requirements) to obtain the total updated training reinforcement benefit.

[0117] Determine the first fusion coefficient of the reference training error under the example learning data according to the first difference between the updated training reinforcement benefit and the target training reinforcement benefit. Assume that the updated training reinforcement benefit is UT, the target training reinforcement benefit is TT, and the first difference is D1 = UT - TT. The first fusion coefficient has a negative correlation with the first difference, which means that when the first difference increases, the first fusion coefficient decreases. To determine this first fusion coefficient, according to the updated training reinforcement benefits and target training reinforcement benefits corresponding to each example learning data in the example training group corresponding to the example learning data, determine the first difference corresponding to each example learning data. For example, there are n example learning data in the example training group, and for each example learning data, the first difference D11, D12,..., D1n is calculated in the above manner. Calculate the mean value of the first differences corresponding to each example learning data to determine the first threshold value matching the example learning data. That is, the first threshold value = (D11 + D12 +... + D1n) / n. When the first difference corresponding to the example learning data is not greater than the first threshold value, determine the first fusion coefficient of the reference training error under the example learning data according to the negative correlation. For example, a negative correlation function relationship can be set, such as the first fusion coefficient = - k * D1 (k is a constant determined according to experience or experiment). When D1 is not greater than the first threshold value, the first fusion coefficient is determined according to this formula.

[0118] Determine the second fusion coefficient of the model training comparison loss under the example learning data. The model training comparison loss is calculated based on the positive supervision training reinforcement benefit and the negative supervision training reinforcement benefit. For the example learning data from 9:00 am to 10:00 am, the positive supervision training reinforcement benefit and the negative supervision training reinforcement benefit are comprehensively calculated by previously inputting the example microgrid operation data into the updated microgrid reactive power optimization model and the candidate microgrid reactive power optimization model, and considering aspects such as reactive power compensation accuracy, voltage stability, and grid loss. Assume the positive supervision training reinforcement benefit is PT, the negative supervision training reinforcement benefit is NT, and the model training comparison loss = PT - NT. Determine the second fusion coefficient of the model training comparison loss under the example learning data according to the second difference between the positive supervision training reinforcement benefit and the negative reactive power configuration supervision data training reinforcement benefit. Assume the positive supervision training reinforcement benefit is PT, the negative reactive power configuration supervision data training reinforcement benefit is NT', and the second difference is D2 = PT - NT'. The second fusion coefficient has a negative correlation with the second difference. Similarly, in the same way as determining the first fusion coefficient, according to the positive supervision training reinforcement benefits and the negative reactive power configuration supervision data training reinforcement benefits corresponding to each example learning data in the example training group corresponding to the example learning data, determine the second differences D21, D22,..., D2n corresponding to each example learning data respectively, calculate the mean of these second differences to obtain the second threshold value matching the example learning data. When the second difference corresponding to the example learning data is not greater than the second threshold value, determine the second fusion coefficient according to the negative correlation, for example, it can be set as the second fusion coefficient = - m * D2 (m is a constant determined by experience or experiment).

[0119] Based on the first fusion coefficient and the second fusion coefficient, perform a fusion calculation on the reference training error and the model training comparison loss of the example learning data to determine the training error corresponding to the example learning data. Assume the reference training error is RE, the model training comparison loss is ML, the first fusion coefficient is α, and the second fusion coefficient is β. Then the training error corresponding to the example learning data = α * RE + β * ML. This training error comprehensively considers the reference training error, the model training comparison loss, and their respective fusion coefficients, and can more comprehensively reflect the state of the example learning data during the training process.

[0120] When the calculation results of the training errors corresponding to each sample learning data no longer continue to decrease, it is determined that the training error in the target training stage no longer continues to decrease. During the entire training process, the training error of each sample learning data is continuously calculated. For example, in a sample training group, there are multiple sample learning data, and as the model parameters are continuously learned, the training error of each sample learning data will change. When the training errors of all sample learning data no longer continue to decrease, it indicates that the entire model has reached a relatively stable state in the target training stage, that is, the training error in the target training stage no longer continues to decrease. This determination basis can effectively determine the termination time of model training and ensure that the obtained model has good performance in the reactive power optimization of the industrial park microgrid.

[0121] In a possible implementation manner, step B122 further includes:

[0122] Step B1224, determining the training reinforcement benefit of the negative reactive power configuration supervision data of the updated microgrid reactive power optimization model corresponding to the sample learning data compared with the candidate microgrid reactive power optimization model.

[0123] Step B1225, determining the second fusion coefficient of the model training comparison loss under the sample learning data according to the second difference between the positive supervision training reinforcement benefit and the training reinforcement benefit of the negative reactive power configuration supervision data. The second fusion coefficient has a negative correlation with the second difference.

[0124] Among them, step B1225 includes:

[0125] Step B1225-1, determining the second difference corresponding to each sample learning data according to the positive supervision training reinforcement benefit and the training reinforcement benefit of the negative reactive power configuration supervision data corresponding to each sample learning data in the sample training group corresponding to the sample learning data.

[0126] Step B1225-2, calculating the mean value of the second differences corresponding to each sample learning data to determine the second threshold value matching the sample learning data.

[0127] Step B1225-3, when the second difference corresponding to the sample learning data is not greater than the second threshold value, determining the second fusion coefficient of the model training comparison loss under the sample learning data according to the negative correlation.

[0128] In this embodiment, taking the operation period of the microgrid from 9:00 to 10:00 in the morning in the sample learning data as an example, within this period, the sample microgrid operation data includes various parameters such as voltage, current, power factor, output power of distributed power sources, and power demand of loads. The sample microgrid operation data is respectively input into the updated microgrid reactive power optimization model and the candidate microgrid reactive power optimization model.

[0129] For the negative reactive power configuration supervision data, corresponding output results are obtained after being input into the two models. Assume that the predicted value of the reactive power compensation amount output by the candidate microgrid reactive power optimization model for the negative reactive power configuration supervision data is Q1', and the predicted value of the reactive power compensation amount output by the updated microgrid reactive power optimization model is Q1''. Considering from the aspect of reactive power compensation accuracy, the training reinforcement benefit (reactive power compensation accuracy) of the negative reactive power configuration supervision data can be obtained by calculating |Q1 - Q1'| - |Q1 - Q1''|, where Q1 is the actual reactive power compensation amount in the negative reactive power configuration supervision data. At the same time, determine the training reinforcement benefit of the negative reactive power configuration supervision data from two aspects that are crucial to the operation of the microgrid, namely voltage stability and grid loss. For voltage stability, calculate the corresponding benefit difference according to the influence of the reactive power configuration output by the two models on the voltage fluctuation range of the microgrid. For example, the voltage fluctuation range of the microgrid under the candidate microgrid reactive power optimization model is ΔV3, and under the updated microgrid reactive power optimization model is ΔV4, then the training reinforcement benefit (voltage stability) of the negative reactive power configuration supervision data in terms of voltage stability is ΔV3 - ΔV4. For grid loss, according to the relationship model between reactive power configuration and grid loss, assume that the grid loss under the candidate microgrid reactive power optimization model is L3, and under the updated microgrid reactive power optimization model is L4, and the training reinforcement benefit (grid loss) of the negative reactive power configuration supervision data is L3 - L4. Combining these aspects (weighted summation according to the pre-determined weights, and the weights are set according to the operation characteristics and optimization objectives of the microgrid), the training reinforcement benefit of the updated microgrid reactive power optimization model corresponding to the sample learning data compared with the candidate microgrid reactive power optimization model for the negative reactive power configuration supervision data is obtained.

[0130] Next, based on the second difference between the reinforcement benefit of active supervision training and the reinforcement benefit of passive reactive power allocation supervision data training, determine the second fusion coefficient of the comparison loss of model training under the sample learning data. The reinforcement benefit of active supervision training is also calculated from aspects such as reactive power compensation accuracy, voltage stability, and grid loss. By inputting the sample microgrid operation data into the updated microgrid reactive power optimization model and the candidate microgrid reactive power optimization model, the output results under the active reactive power allocation supervision data are compared. Assume that the reinforcement benefit of active supervision training is PT, the reinforcement benefit of passive reactive power allocation supervision data training is NT, and the second difference is D2 = PT - NT. Here, the second fusion coefficient has a negative correlation with the second difference, meaning that the larger the second difference, the smaller the second fusion coefficient.

[0131] To determine this second fusion coefficient, based on the reinforcement benefit of active supervision training and the reinforcement benefit of passive reactive power allocation supervision data training corresponding to each sample learning data in the sample training group corresponding to the sample learning data, determine the second difference corresponding to each sample learning data. For example, in the sample training group, there are n sample learning data. For each sample learning data, calculate the reinforcement benefit of active supervision training and the reinforcement benefit of passive reactive power allocation supervision data training in the above manner, and obtain the corresponding second differences D21, D22, …, D2n.

[0132] Then, calculate the mean of the second differences corresponding to each sample learning data to determine the second threshold value matching the sample learning data. That is, calculate the second threshold value = (D21 + D22 + … + D2n) / n. When the second difference corresponding to the sample learning data is not greater than the second threshold value, determine the second fusion coefficient of the comparison loss of model training under the sample learning data according to the negative correlation. For example, a functional relationship can be set to reflect this negative correlation. Assume that the second fusion coefficient is β, β = - m * D2 (m is a constant determined according to experience or experiment). When the second difference D2 corresponding to the sample learning data is not greater than the second threshold value, determine the second fusion coefficient of the comparison loss of model training under the sample learning data according to this formula. This way of determining the second fusion coefficient comprehensively considers the situations of each sample learning data in the sample training group, making the determination of the second fusion coefficient more reasonable and accurate. Furthermore, when calculating the training error corresponding to the sample learning data later, it can more comprehensively reflect the training state of the model, which helps to improve the accuracy and reliability of the microgrid reactive power optimization model to better meet the operation requirements of the industrial park microgrid.

[0133] In a possible implementation manner, step S110 includes:

[0134] Step S111, obtain the reactive power optimization requirements of the microgrid and multiple sample microgrid operation data.

[0135] Step S112: For each piece of the sample microgrid operation data, determine the demand index data associated with the reactive power optimization demand of the microgrid. The demand index data includes the sample microgrid operation data.

[0136] Step S113: Load the demand index data into a pre-trained deep learning network, and use the output of the deep learning network as the negative reactive power configuration supervision data for the sample microgrid operation data.

[0137] Step S114: Generate the positive reactive power configuration supervision data for the sample microgrid operation data according to the expert optimization instructions generated based on the negative reactive power configuration supervision data for the sample microgrid operation data.

[0138] Step S115: Configure the sample learning data including the sample microgrid operation data, the positive reactive power configuration supervision data for the sample microgrid operation data, and the negative reactive power configuration supervision data for the sample microgrid operation data, and generate a training data sequence including multiple pieces of sample learning data.

[0139] In this embodiment, the reactive power optimization demand of the microgrid is determined based on the overall operation objectives of the industrial park microgrid, and these objectives include but are not limited to maintaining voltage stability, reducing grid losses, improving power quality, etc. For example, in order to meet the requirements of various industrial equipment in the industrial park for power quality, it is necessary to control the voltage fluctuation range within a certain value, which is a specific reactive power optimization demand. At the same time, collect multiple pieces of sample microgrid operation data, which reflect the actual operation status of the microgrid at different time periods. Taking different time periods within a day as an example, such as from 9:00 am to 10:00 am, at this time, industrial production is in a normal production state, most equipment is operating at full load, the illumination condition of the photovoltaic panels is of medium intensity, and the corresponding sample microgrid operation data includes the voltage value of the microgrid being 380V, the current being 50A, the power factor being 0.85, the output power of each distributed power source (such as solar photovoltaic panels, small wind turbines, etc.), and the power demands of different types of loads (industrial electrical equipment, office area electrical equipment, etc.).

[0140] For each sample microgrid operation data, determine the demand index data associated with the reactive power optimization requirements of the microgrid. The demand index data includes the sample microgrid operation data itself and other derivative data related to the reactive power optimization requirements. For example, in addition to the direct operation data such as voltage, current, and power factor mentioned above, it may also include data such as apparent power and reactive power calculated based on voltage and current. These demand index data can more comprehensively reflect the relationship between the operation state of the microgrid and the reactive power optimization requirements. Taking the sample microgrid operation data from 9 am to 10 am as an example, the calculated data such as apparent power and reactive power are closely related to the reactive power optimization requirements such as pre-set voltage stability and grid loss, and together constitute the demand index data.

[0141] Load the demand index data into a pre-trained deep learning network, and use the output of the deep learning network as the negative reactive power configuration supervision data for the sample microgrid operation data. This pre-trained deep learning network is trained based on a large amount of microgrid operation data and related reactive power configuration cases. When the demand index data corresponding to the sample microgrid operation data from 9 am to 10 am is loaded into this deep learning network, the network outputs a reactive power configuration scheme as the negative reactive power configuration supervision data according to its internal algorithms and model structure. For example, based on the input data, the deep learning network may determine that the reactive power compensation amount to be set on a certain reactive power compensation device is Q1, and the reactive power configuration scheme corresponding to this Q1 is the negative reactive power configuration supervision data for this sample microgrid operation data. This negative reactive power configuration supervision data is a preliminary reactive power configuration suggestion obtained based on the deep learning network's learning of existing data and rules.

[0142] Generate the positive reactive power configuration supervision data for the sample microgrid operation data according to the expert optimization instructions generated based on the negative reactive power configuration supervision data for the sample microgrid operation data. The expert optimizes the negative reactive power configuration supervision data based on their rich experience and deeper understanding of the industrial park microgrid system. The expert will consider more complex factors in actual operation, such as the sensitivity of specific industrial equipment to reactive power and the output characteristics of distributed power sources under different seasons and weather conditions. Taking the reactive power compensation amount Q1 in the negative reactive power configuration supervision data as an example, the expert may believe that in the current operation state, considering the upcoming peak electricity consumption period and the special requirements of a key industrial equipment for voltage stability, it is more appropriate to adjust the reactive power compensation amount to Q2, and the reactive power configuration scheme corresponding to this Q2 is used as the positive reactive power configuration supervision data. The expert optimization instructions are adjustments made based on the negative reactive power configuration supervision data, taking into account various actual factors, which are more conducive to meeting the reactive power optimization requirements of the microgrid.

[0143] Finally, configure example learning data including example microgrid operation data, positive reactive power configuration supervision data of the example microgrid operation data, and negative reactive power configuration supervision data of the example microgrid operation data, and generate a training data sequence including multiple example learning data. For the example from 9 am to 10 am, combine the example microgrid operation data, the corresponding positive reactive power configuration supervision data Q2, and the negative reactive power configuration supervision data Q1 in this period to form an example learning data. In the same way, process the example microgrid operation data in other different periods to obtain multiple example learning data. These example learning data together constitute the training data sequence, which will be used for the subsequent training of the microgrid reactive power optimization model, providing rich learning samples for the model, enabling it to better adapt to various operating conditions of the industrial park microgrid, and thus improving the accuracy and effectiveness of reactive power optimization.

[0144] In a possible implementation manner, the process of obtaining a candidate microgrid reactive power optimization model includes:

[0145] Obtain an initialized neural network model.

[0146] According to the training data sequence, configure a basic training data sequence including multiple basic example learning data. Each of the basic example learning data includes example microgrid operation data, and positive reactive power configuration supervision data of the example microgrid operation data or negative reactive power configuration supervision data of the example microgrid operation data.

[0147] Iteratively update the neural network model according to the basic training data sequence to generate a candidate microgrid reactive power optimization model.

[0148] In this embodiment, the initialized neural network model is the basis for constructing the microgrid reactive power optimization model. For example, select a neural network model with a multi-layer perceptron (MLP) structure, which has an input layer, several hidden layers, and an output layer. The number of nodes in the input layer is determined according to the characteristics of the example microgrid operation data. Features such as voltage, current, power factor, distributed power source output power, and load power demand in the industrial park microgrid can all be used as inputs to the input layer. The number of neurons and the number of layers in the hidden layer are set according to the complexity requirements of the model. For example, 3 hidden layers can be set, and each layer has a different number of neurons to achieve complex non-linear mapping of the input data. The output of the output layer corresponds to the relevant results of reactive power configuration, such as reactive power compensation amount, etc. The initial weights, biases, and other parameters of this initialized neural network model may be randomly set. It is just a model with a basic structure but not yet trained for microgrid reactive power optimization.

[0149] According to the training data sequence, configure a basic training data sequence containing multiple basic example learning data. Each basic example learning data includes example microgrid operation data, and positive reactive power configuration supervision data of the example microgrid operation data or negative reactive power configuration supervision data of the example microgrid operation data. Extract and organize from the previously obtained training data sequence. Taking the example of the industrial park microgrid from 9 am to 10 am as an example, its example microgrid operation data includes information such as a voltage of 380V, a current of 50A, and a power factor of 0.85. The basic example learning data can only include this example microgrid operation data and the corresponding negative reactive power configuration supervision data. For example, the reactive power compensation amount of the reactive power compensation device in the negative reactive power configuration supervision data is Q1. Or it can also be the example microgrid operation data and positive reactive power configuration supervision data, such as the reactive power compensation amount in the positive reactive power configuration supervision data is Q2. In this way, multiple examples in the training data sequence are processed to form multiple basic example learning data, and these basic example learning data together constitute the basic training data sequence.

[0150] Iteratively update the neural network model according to the basic training data sequence to generate a candidate microgrid reactive power optimization model. Input the basic example learning data in the basic training data sequence into the initialized neural network model one by one. For each basic example learning data, the neural network model performs forward propagation calculation according to the input example microgrid operation data to obtain a preliminary reactive power configuration output result. Then, according to the difference between the output result and the corresponding positive reactive power configuration supervision data or negative reactive power configuration supervision data, calculate the loss function. For example, use the mean square error (MSE) as the loss function. If the input is the example microgrid operation data and the negative reactive power configuration supervision data, and the reactive power compensation amount in the negative reactive power configuration supervision data is Q1, and the predicted value of the reactive power compensation amount output by the model is Q1', then the value of the loss function is (Q1 - Q1')². Then, through the backpropagation algorithm, adjust the parameters such as the weights and biases of the neural network model according to the loss function to reduce the value of the loss function. This process is repeated continuously. As more basic example learning data are input into the model and iteratively updated, the neural network model gradually learns the relationship pattern between the example microgrid operation data and the reactive power configuration. After multiple rounds of iteration, when the value of the loss function reaches a stable and acceptable range, the neural network model at this time becomes the candidate microgrid reactive power optimization model. This candidate microgrid reactive power optimization model already has a certain ability to make preliminary optimization decisions on the reactive power configuration of the microgrid, laying a foundation for subsequent further training and optimization.

[0151] Figure 2The figure shows the hardware structure diagram of the deep learning-based microgrid reactive power optimization system 100 provided by the embodiments of the present invention for implementing the above-mentioned deep learning-based microgrid reactive power optimization method, as Figure 2 shown, the deep learning-based microgrid reactive power optimization system 100 may include a processor 110, a machine-readable storage medium 120, a bus 130, and a communication unit 140.

[0152] The machine-readable storage medium 120 may store data and / or instructions. In some embodiments, the machine-readable storage medium 120 may store data obtained from an external terminal. In some embodiments, the machine-readable storage medium 120 may store the data and / or instructions used by the deep learning-based microgrid reactive power optimization system 100 to execute or use to complete the exemplary methods described in the present invention.

[0153] In a specific implementation process, one or more processors 110 execute the computer-executable instructions stored in the machine-readable storage medium 120, so that the processor 110 can execute the deep learning-based microgrid reactive power optimization method of the above method embodiment. The processor 110, the machine-readable storage medium 120, and the communication unit 140 are connected through the bus 130, and the processor 110 can be used to control the transceiver actions of the communication unit 140.

[0154] For the specific implementation process of the processor 110, reference may be made to the respective method embodiments executed by the above-mentioned deep learning-based microgrid reactive power optimization system 100. Their implementation principles and technical effects are similar, and will not be elaborated here in this embodiment.

[0155] In addition, an embodiment of the present invention also provides a readable storage medium, in which computer-executable instructions are preset. When the processor executes the computer-executable instructions, the above-mentioned deep learning-based microgrid reactive power optimization method is implemented.

[0156] It should be noted that, in order to simplify the description of the present invention disclosure and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, sometimes multiple features are merged into one embodiment, drawing, or description thereof.

Claims

1. A microgrid reactive power optimization method based on deep learning, characterized in that: The method comprises: Obtain a candidate microgrid reactive optimization model and a training data sequence containing a plurality of sample learning data; the sample learning data contains sample microgrid operation data, passive reactive configuration supervision data of the sample microgrid operation data, and active reactive configuration supervision data of the sample microgrid operation data, wherein the step of determining the passive reactive configuration supervision data comprises: loading the sample microgrid operation data as demand indicator data into a pre-trained deep learning network, the deep learning network outputting a reactive configuration scheme as passive reactive configuration supervision data according to the input microgrid operation data and the pre-learned rules and algorithms; and optimizing the passive reactive configuration supervision data according to the expert optimization instructions generated for the passive reactive configuration supervision data to obtain the active reactive configuration supervision data; For each sample learning data in the training data sequence, when model parameter learning is performed according to the sample learning data in a target training phase, determining a target training reinforcement benefit of the positive reactive configuration supervision data in the sample learning data compared to the negative reactive configuration supervision data; Determine a first cycle goodness of fit value of the updated microgrid reactive optimization model corresponding to the sample learning data to generate active reactive configuration supervision data, and a second cycle goodness of fit value of the updated microgrid reactive optimization model corresponding to the sample learning data to generate passive reactive configuration supervision data; the updated microgrid reactive optimization model corresponding to the sample learning data is generated by model parameter learning based on the previous sample training group of the sample training group corresponding to the sample learning data; in the process of the first model parameter learning, the deep learning target of the first sample training group is the candidate microgrid reactive optimization model, and the first cycle goodness of fit value and the second cycle goodness of fit value are calculated according to the difference between the predicted value and the actual value; The first cycle fit goodness of fit value and the second cycle fit goodness of fit value are fused and calculated to determine the cycle training reinforcement benefit of the sample learning data, and when it is determined that the training error of the target training stage no longer continues to decrease according to the cycle training reinforcement benefit and the target training reinforcement benefit corresponding to each of the sample learning data, the model parameter learning is terminated to generate the target microgrid reactive optimization model corresponding to the candidate microgrid reactive optimization model; the fusion coefficient of the first cycle fit goodness of fit value is less than the fusion coefficient of the second cycle fit goodness of fit value, wherein the cycle training reinforcement benefit is calculated by the formula: cycle training reinforcement benefit = α*F1+β*F2, wherein F1 is the first cycle fit goodness of fit value, F2 is the second cycle fit goodness of fit value, α and β are the fusion coefficients of the first cycle fit goodness of fit value and the second cycle fit goodness of fit value, respectively, α<β and α+β=1; The target microgrid operation data to be analyzed is obtained, the target microgrid operation data is loaded into the target microgrid reactive power optimization model, and reactive power optimization configuration decision data of the target microgrid operation data is generated.

2. The microgrid reactive power optimization method based on deep learning according to claim 1, characterized in that: The determining of the first cycle goodness of fit value of the updated microgrid reactive optimization model corresponding to the sample learning data to generate active reactive configuration supervision data includes: Acquire multiple candidate active reactive configuration supervision data of the sample learning data; For each of the candidate active reactive configuration supervision data, determining a first cycle goodness of fit value of the updated microgrid reactive optimization model corresponding to the sample learning data to generate the candidate active reactive configuration supervision data; The first cycle goodness of fit values ​​corresponding to each of the candidate active reactive configuration supervision data are averaged to generate the first cycle goodness of fit value of the active reactive configuration supervision data generated by the updated microgrid reactive optimization model corresponding to the sample learning data.

3. The microgrid reactive power optimization method based on deep learning according to claim 1, characterized in that: The method further comprises: Determine the positive supervision training reinforcement benefit of the updated microgrid reactive optimization model corresponding to the sample learning data compared to the candidate microgrid reactive optimization model, and the negative supervision training reinforcement benefit of the updated microgrid reactive optimization model corresponding to the sample learning data compared to the candidate microgrid reactive optimization model; Determining the model training comparison loss corresponding to the sample learning data according to the positive supervision training reinforcement benefit and the negative supervision training reinforcement benefit; When it is determined that the training error of the target training stage no longer continues to decrease according to the cyclic training reinforcement benefit and the target training reinforcement benefit respectively corresponding to each of the sample learning data, the model parameter learning is terminated, and the target microgrid reactive power optimization model corresponding to the candidate microgrid reactive power optimization model is generated, including: When it is determined that the training error of the target training stage no longer continues to decrease based on the cyclic training enhancement benefit, the target training enhancement benefit and the model training comparison loss corresponding to each of the sample learning data, the model parameter learning is terminated and the target microgrid reactive optimization model corresponding to the candidate microgrid reactive optimization model is generated.

4. The microgrid reactive power optimization method based on deep learning according to claim 3 is characterized in that: Determining the negative supervision training reinforcement benefit of the updated microgrid reactive power optimization model corresponding to the sample learning data compared with the candidate microgrid reactive power optimization model, including: Determine the training enhancement benefits of the passive reactive configuration supervision data corresponding to each of the candidate microgrid reactive optimization models in the sample training group corresponding to the sample learning data and under each of the sample learning data; The mean value of each of the passive reactive configuration supervision data training enhancement benefits is calculated to determine the passive supervision training enhancement benefit of the updated microgrid reactive optimization model corresponding to the sample learning data compared with the candidate microgrid reactive optimization model.

5. The microgrid reactive power optimization method based on deep learning according to claim 3 is characterized in that: The method further comprises: For each of the sample learning data, determining a reference training error corresponding to the sample learning data according to the cyclic training reinforcement benefit and the target training reinforcement benefit of the sample learning data; Determine a first fusion coefficient of a reference training error under the sample learning data and a second fusion coefficient of a model training comparison loss under the sample learning data; According to the first fusion coefficient and the second fusion coefficient, a reference training error of the sample learning data and a model training comparison loss are fused and calculated to determine a training error corresponding to the sample learning data; When the calculation results of the training errors corresponding to the sample learning data no longer continue to decrease, it is determined that the training error of the target training stage no longer continues to decrease.

6. The microgrid reactive power optimization method based on deep learning according to claim 5, characterized in that: The determining of a first fusion coefficient of a reference training error under the sample learning data includes: Determine the update training reinforcement benefit of the active reactive configuration supervision data under the sample learning data compared with the passive reactive configuration supervision data after completing the model parameter learning; Determining a first fusion coefficient of the reference training error under the sample learning data according to a first difference between the updated training reinforcement benefit and the target training reinforcement benefit; the first fusion coefficient and the first difference are negatively correlated; The step of determining a first fusion coefficient of a reference training error under the sample learning data according to a first difference between the updated training enhancement benefit and the target training enhancement benefit includes: Determine first differences corresponding to each sample learning data according to the updated training reinforcement benefits and the target training reinforcement benefits corresponding to each sample learning data in the sample training group corresponding to the sample learning data; Calculate the mean of the first difference values ​​respectively corresponding to the sample learning data to determine a first threshold value matching the sample learning data; When the first difference corresponding to the sample learning data is not greater than the first threshold value, a first fusion coefficient of the reference training error under the sample learning data is determined according to the negative correlation relationship.

7. The microgrid reactive power optimization method based on deep learning according to claim 5, characterized in that: Determining a second fusion coefficient of the model training comparison loss under the sample learning data includes: Determine the training enhancement benefit of the passive reactive configuration supervision data of the updated microgrid reactive optimization model corresponding to the sample learning data compared to the candidate microgrid reactive optimization model; Determine a second fusion coefficient of the model training comparison loss under the sample learning data according to a second difference between the active supervision training enhancement benefit and the passive reactive configuration supervision data training enhancement benefit; the second fusion coefficient and the second difference are negatively correlated; Wherein, determining the second fusion coefficient of the model training comparison loss under the sample learning data according to the second difference between the active supervision training enhancement benefit and the passive reactive configuration supervision data training enhancement benefit includes: Determine the second difference value corresponding to each sample learning data according to the positive supervision training reinforcement benefit and the negative reactive configuration supervision data training reinforcement benefit corresponding to each sample learning data in the sample training group corresponding to the sample learning data; Calculating the mean of the second difference values ​​respectively corresponding to the sample learning data to determine a second threshold value matching the sample learning data; When the second difference corresponding to the sample learning data is not greater than the second threshold value, a second fusion coefficient of the model training comparison loss under the sample learning data is determined according to the negative correlation relationship.

8. The microgrid reactive power optimization method based on deep learning according to any one of claims 1 to 7, characterized in that: The process of obtaining a training data sequence containing multiple sample learning data includes: Obtain microgrid reactive power optimization requirements and multiple sample microgrid operation data; For each of the sample microgrid operation data, determining demand index data associated with the microgrid reactive power optimization demand; the demand index data includes the sample microgrid operation data; Loading the demand index data into a pre-trained deep learning network, and using the generation of the deep learning network as passive reactive configuration supervision data of the sample microgrid operation data; Generate active reactive configuration supervision data for the sample microgrid operation data according to the expert optimization instruction generated for the passive reactive configuration supervision data for the sample microgrid operation data; The sample learning data including the sample microgrid operation data, the active reactive configuration supervision data of the sample microgrid operation data and the passive reactive configuration supervision data of the sample microgrid operation data are configured to generate a training data sequence including a plurality of sample learning data.

9. The microgrid reactive power optimization method based on deep learning according to any one of claims 1 to 7, characterized in that: The process of obtaining a candidate microgrid reactive power optimization model includes: Get the initialized neural network model; According to the training data sequence, a basic training data sequence including a plurality of basic sample learning data is configured; each of the basic sample learning data includes sample microgrid operation data, and active reactive configuration supervision data of the sample microgrid operation data or passive reactive configuration supervision data of the sample microgrid operation data; The neural network model is iteratively updated according to the basic training data sequence to generate a candidate microgrid reactive power optimization model.

10. A microgrid reactive power optimization system based on deep learning, characterized in that: The deep learning-based microgrid reactive power optimization system includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the deep learning-based microgrid reactive power optimization method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Power grid reactive active prediction and control technology based on situation awareness

    CN113964885A

  • Reactive voltage optimization method and device based on reinforcement learning, equipment and medium

    CN115833147A