Power system unit scheduling optimization method and device, electronic equipment and storage medium
Through deep learning and extreme random forest algorithms, the difficulty of power system unit scheduling is predicted, and the solver parameters are dynamically adjusted, which solves the problem of poor solution efficiency and effect in the existing technology, and realizes intelligent optimization and efficient solution of power system unit scheduling.
Patent Information
- Application Number
- CN202510241992.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-03
AI Technical Summary
When dealing with the combination of power system units, the prior art lacks the ability to adapt to the complexity of the problem, resulting in significant differences in solution efficiency and effect. Especially when facing the changes in the access to new energy and the balance between power supply and demand, the challenges are even more severe.
Through deep learning of historical data characteristics, the extreme random forest algorithm is used to predict the difficulty of clearing data, and the solver parameters are dynamically adjusted to achieve intelligent optimization of power system unit scheduling.
It significantly improves the solution efficiency, ensures the accuracy and feasibility of the scheduling results, and can effectively deal with the cleaning problems of power systems of different types and scales.
Smart Images

Figure CN120090277A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system dispatching, and particularly relates to a method for optimizing unit dispatching in a power system, a device for optimizing unit dispatching in a power system, an electronic device, and a computer-readable storage medium. Background Art
[0002] In the operation of a power system, the unit commitment problem involves determining the start-stop states and output powers of each unit according to factors such as load demand, unit characteristics, operation constraints, and economic costs within a given time period. In related technologies, a mixed-integer linear programming model is generally used to model the unit commitment problem, and a mathematical solver is used for solving. However, when dealing with such problems, the solver often adopts a fixed solution strategy and parameter configuration, lacking the ability to adaptively adjust according to the complexity of the problem. As a result, when the solver faces clearing problems of different types and scales, there are significant differences in its solution efficiency and effect. Summary of the Invention
[0003] The present invention is made based on the inventor's discovery and recognition of the following facts and problems:
[0004] With the continuous expansion of the scale of the power system, especially the rapid development and large-scale grid connection of intermittent new energy (such as wind power, solar energy, etc.) in recent years, the complexity of the unit commitment problem has shown an explosive growth. On the one hand, the intermittency and uncertainty of new energy power generation have brought unprecedented challenges to the solution mode based on fixed solution strategies and parameter configurations; on the other hand, the access of new energy has also changed the original power supply-demand balance, making the solution space of the unit commitment problem wider and the constraint conditions more complex and changeable.
[0005] In related technologies, when using a solver to solve the unit commitment problem, a fixed solution strategy and parameter configuration are often adopted, such as fixed iteration times, fixed error accuracy, etc. The selection of these fixed parameters is usually based on experience or default values, lacking the ability to adaptively adjust according to the specific complexity of the problem. However, when the solver faces clearing problems of different types (such as a power system with a large amount of new energy or a high proportion of traditional thermal power units) and different scales (such as provincial power grids and national power grids), there are often significant differences in its solution efficiency and effect. For example, in some cases, the solver may get stuck in a long solution process due to an overly conservative strategy, resulting in a large waste of computing resources; in other cases, it may not be able to find a high-quality solution within the specified time due to an overly aggressive strategy, and may even directly threaten the stable operation of the power system due to solution failure.
[0006] To this end, the present invention provides a method for optimizing unit scheduling in a power system, a device for optimizing unit scheduling in a power system, an electronic device, and a computer-readable storage medium, which can realize intelligent optimization of unit scheduling in a power system, not only effectively ensure the accuracy and feasibility of the scheduling result, but also improve the solution efficiency.
[0007] The method for optimizing unit scheduling in a power system provided by the present invention includes the following steps:
[0008] Obtain the system clearing historical data and the clearing data to be processed;
[0009] Extract features and label the system clearing historical data to establish a sample data set with label information;
[0010] Based on the extreme random forest algorithm, process the sample data in the sample data set to establish an extreme random forest prediction model;
[0011] Substitute the clearing data to be processed into the extreme random forest prediction model to obtain label information matching the clearing data;
[0012] Based on the label information and a preset parameter matching model, determine the solver parameters matching the clearing data;
[0013] Substitute the clearing data and the solver parameters into a preset unit scheduling solver to obtain a unit scheduling plan matching the clearing data.
[0014] In summary, the method for optimizing unit scheduling in a power system provided by the present invention realizes intelligent optimization of unit scheduling in a power system through steps such as deep learning of historical data features, predicting the difficulty of clearing data using the extreme random forest algorithm, and dynamically adjusting solver parameters. This method can not only significantly improve the solution efficiency, but also effectively ensure the accuracy and feasibility of the scheduling result.
[0015] In some embodiments, the step of processing the sample data in the sample data set based on the extreme random forest algorithm to establish an extreme random forest prediction model includes:
[0016] According to the label information of the sample data, divide the sample data in the sample data set into multiple subsets;
[0017] Based on the extreme random forest algorithm, process each subset to establish an extreme random forest decision tree. There are at least two extreme random forest decision trees, and each extreme random forest decision tree corresponds to a subset;
[0018] Construct an extreme random forest prediction model according to all the extreme random forest decision trees.
[0019] In some embodiments, after the step of constructing the extremely randomized forest prediction model, the power system unit scheduling optimization method further includes the steps of:
[0020] According to the multi-fold verification algorithm, the divided subsets are divided into a training subset and a verification subset, and there are at least two training subsets;
[0021] According to the sample data in the sample data set, true positive samples, false positive samples, false negative samples, and true negative samples are established;
[0022] The true positive samples, false positive samples, false negative samples, and true negative samples are supplemented to the verification subset to establish a verification test set;
[0023] The sample data in the verification test set is substituted into the extremely randomized forest prediction model, and model evaluation parameters are calculated. The model evaluation parameters include model accuracy, macro-average precision, macro-average recall rate, macro-average F1 score, and macro-average area under the receiver operating characteristic curve. The macro-average F1 score is the macro-average value after the harmonic mean of the precision rate and the recall rate;
[0024] The reliability score of the extremely randomized forest prediction model is calculated by using the model evaluation parameters and a preset overall reliability evaluation model;
[0025] If the reliability score and the model evaluation parameters meet the preset conditions, it is considered that the extremely randomized forest prediction model runs reliably. The preset conditions include that at least one of the reliability score is greater than or equal to the score threshold, the macro-average recall rate is greater than or equal to the recall rate threshold, and the macro-average area under the receiver operating characteristic curve is greater than or equal to the area threshold under the receiver operating characteristic curve.
[0026] In some embodiments, the step of dividing the sample data in the sample data set into multiple subsets according to the label information of the sample data includes:
[0027] The sample data in the sample data set is classified according to the label information of the sample data;
[0028] Based on the stratified sampling algorithm, the classified sample data is divided into multiple subsets, and there are at least two subsets.
[0029] In some embodiments, the system clearing historical data includes operation data and the corresponding solution time of the scheme. The step of performing feature extraction and label annotation on the system clearing historical data to establish a sample data set with label information includes:
[0030] Extract features from the operation data in the historical data cleared by the system, and preprocess the extracted features. The preprocessing includes at least filling in missing values, removing outliers, and data standardization processing;
[0031] According to the solution time corresponding to each system cleared historical data and a preset label annotation model, determine the label information matching the system cleared historical data. The label annotation model includes multiple label information and multiple time nodes, and the multiple time nodes gradually increase, and the label information corresponds to the time nodes one by one;
[0032] According to the operation data, the solution time of the plan, and the label information, establish a sample data set with label information.
[0033] In some embodiments, the label information is divided into a first difficulty label, a second difficulty label, and a third difficulty label, and the time nodes are divided into a first time threshold and a second time threshold; the step of determining the label information matching the system cleared historical data according to the solution time corresponding to each system cleared historical data and a preset label annotation model includes:
[0034] Judge whether the solution time corresponding to the system cleared historical data is less than the first time threshold;
[0035] If the solution time is less than the first time threshold, then match the system cleared historical data with the first difficulty label;
[0036] If the solution time is greater than or equal to the first time threshold, then judge whether the solution time is less than or equal to the second time threshold, and the second time threshold is greater than the first time threshold;
[0037] If the solution time is greater than or equal to the first time threshold and less than or equal to the second time threshold, then match the system cleared historical data with the second difficulty label;
[0038] If the solution time is greater than or equal to the second time threshold, then match the system cleared historical data with the third difficulty label.
[0039] In some embodiments, the unit scheduling solver includes an objective function and constraint conditions with the goal of minimizing the sum of start-stop costs and operating costs. The constraint conditions include at least one of a power demand constraint condition, a unit start-stop constraint condition, a unit power output constraint condition, a start-stop change constraint condition, a running ramp constraint condition, and a minimum running time constraint condition;
[0040] And / or, the solver parameters at least include the maximum number of iterations, convergence accuracy, allowable relative error, number of parallel processing threads, heuristic algorithm switch, and setting of the initial relaxation variable.
[0041] In addition, the power system unit scheduling optimization device provided by an embodiment of the present invention includes:
[0042] An acquisition module, which is used to acquire the system clearing historical data and the clearing data to be processed;
[0043] A sample annotation module, which is used to perform feature extraction and label annotation on the system clearing historical data to establish a sample data set with label information;
[0044] A model construction module, which is used to process the sample data in the sample data set according to the extremely randomized forest algorithm to establish an extremely randomized forest prediction model;
[0045] A label matching module, which is used to substitute the clearing data to be processed into the extremely randomized forest prediction model to obtain the label information matching the clearing data;
[0046] A parameter matching module, in which a parameter matching model is provided. The parameter matching module is used to determine the solver parameters matching the clearing data according to the label information of the clearing data;
[0047] A scheme determination module, in which a unit scheduling solver is provided. The scheme determination module is used to process the clearing data and the solver parameters by using the unit scheduling solver to obtain a unit scheduling scheme matching the clearing data.
[0048] In addition, the electronic device provided by an embodiment of the present invention includes a processor and a memory. The memory stores machine-readable instructions executable by the processor. When the machine-readable instructions are executed by the processor, the steps in the power system unit scheduling optimization method provided by any one of the above embodiments are executed.
[0049] In addition, the computer-readable storage medium provided by an embodiment of the present invention stores machine-readable instructions. When the storage medium runs on an electronic device, the machine-readable instructions are used to cause the electronic device to execute the steps in the power system unit scheduling optimization method provided by any one of the above embodiments. Description of the Drawings
[0050] The drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the drawings:
[0051] Figure 1 It is a schematic flow chart of the power system unit scheduling optimization method provided by an embodiment of the present invention.
[0052] Figure 2 It is a partial feature list of the system clearing historical data in the power system unit scheduling optimization method provided by an embodiment of the present invention.
[0053] Figure 3 It is a schematic structural diagram of the power system unit scheduling optimization device provided by an embodiment of the present invention.
[0054] Figure 4 It is a schematic hardware structure diagram of the electronic device provided by an embodiment of the present invention.
[0055] Reference numerals:
[0056] 100, power system unit scheduling optimization device; 110, acquisition module; 120, sample annotation module; 130, model construction module; 140, label matching module; 150, parameter matching module; 160, solution determination module;
[0057] 210, processor; 220, memory; 230, communication interface; 240, communication bus. Detailed implementation manners
[0058] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as a limitation to the present invention.
[0059] Refer to Figure 1 , which is a schematic flow chart of the power system unit scheduling optimization method provided by an embodiment of the present invention. The power system unit scheduling optimization method includes the following steps:
[0060] S10, acquire system clearing historical data and clearing data to be processed.
[0061] Among them, the system clearing historical data contains the historical records of the power system under various operating conditions in the past period. The clearing data to be processed reflects the current state and demand of the power system and is the target object that needs to be optimized for scheduling.
[0062] S20, perform feature extraction and label annotation on the system clearing historical data to establish a sample data set with label information.
[0063] Among them, feature extraction is performed based on the historical system clearing data, and the extracted features are labeled with tags, such as tag information representing difficulty levels like simple questions, medium difficulty, difficult questions, etc., and a sample data set containing the tag information is constructed accordingly.
[0064] S30. Based on the extremely randomized forest algorithm, process the sample data in the sample data set to establish an extremely randomized forest prediction model.
[0065] S40. Substitute the clearing data to be processed into the extremely randomized forest prediction model to obtain the tag information matching the clearing data.
[0066] S50. Based on the tag information and a preset parameter matching model, determine the solver parameters matching the clearing data.
[0067] S60. Substitute the clearing data and the solver parameters into a preset unit scheduling solver to obtain a unit scheduling plan matching the clearing data.
[0068] In the implementation process of the power system unit scheduling optimization method provided in the above embodiment, first, obtain the historical system clearing data in the power system and the clearing data to be processed currently or in the future; extract the features in the historical system clearing data, and label the extracted features to establish a sample data set containing tag information.
[0069] Subsequently, use the extremely randomized forest algorithm to train the sample data set with labeled tags. By randomly selecting features and constructing decision trees, capture the complex non-linear relationship between the features and the tag information, and establish an extremely randomized forest prediction model; after completing the model training, the newly collected clearing data to be processed can be input into the established extremely randomized forest prediction model. The model will quickly predict and output the corresponding tag information according to the features of the input data, that is, the difficulty level of the clearing data.
[0070] Then, according to the preset parameter matching model, dynamically adjust the parameters of the unit scheduling solver. This step sets a set of optimal or sub-optimal solver parameters for each difficulty level based on historical experience and experimental results to ensure that the solving process can be completed efficiently and the quality of the solution can be maintained.
[0071] Finally, combine the solver with optimized parameters with the clearing data to be processed, start the unit scheduling solver, and the solver will output an optimal or approximately optimal unit scheduling plan for the current power system state, thereby realizing the real-time optimization of the operation of the power system.
[0072] In summary, the power system unit scheduling optimization method provided by the present invention realizes the intelligent optimization of power system unit scheduling through steps such as deep learning historical data features, using the extreme random forest algorithm to predict the difficulty of clearing data, and dynamically adjusting solver parameters. It can not only effectively ensure the accuracy and feasibility of the scheduling results, but also improve the solving efficiency.
[0073] In some embodiments, the system clearing historical data includes operation data and the corresponding solution time of the scheme. The steps of extracting features and labeling tags from the system clearing historical data to establish a sample data set with tag information include:
[0074] Preprocess the system clearing historical data, and the preprocessing at least includes filling missing values, removing outliers, and data standardization processing;
[0075] According to the solution time of the scheme corresponding to each system clearing historical data and a preset tag labeling model, determine the tag information matching the system clearing historical data. The tag labeling model includes multiple tag information and multiple time nodes, and the multiple time nodes gradually increase, and the tag information corresponds to the time nodes one by one;
[0076] Establish a sample data set with tag information according to the operation data, the solution time of the scheme, and the tag information.
[0077] Specifically, in the preprocessing stage, the missing values and outliers can be processed first to ensure the consistency and accuracy of the data. The interpolation method can be used to fill the missing values. Of course, in some embodiments, other interpolation methods, such as linear interpolation method, can also be used to fill the missing values to ensure the integrity of the data.
[0078] Assume that in the jth sample, the value of feature x i is x i,j missing, then the missing value can be filled by the mean value of this feature in other samples, which can be expressed by the following formula:
[0079]
[0080] In the formula, x i,k is the value of feature i in the kth sample, and n represents the number of all non-missing value samples.
[0081] For the treatment of outliers, the interquartile range method (Interquartile Range, IQR) can be used for detection and removal. Assume that the first quartile of feature x i is Q1, and the third quartile is Q3, then the interquartile range of this feature is IQR = Q3 - Q1. Among them, the treatment of outliers can include the following steps:
[0082] According to the IQR principle, judge whether x i,j exceeds the preset range, and the preset range is set as x i,j < Q1 - 1.5·IQR or x i,j < Q3 + 1.5·IQR;
[0083] If x i,j exceeds the preset range, then x i,j is regarded as an outlier and is excluded;
[0084] If x i,j does not exceed the preset range, then the x i,j is retained.
[0085] Data standardization processing can eliminate the influence of different dimensions on model training and ensure that each feature is processed on the same scale. The steps of the data standardization processing can be expressed by the following formula:
[0086]
[0087] In the formula, x' i,j is the standardized feature value, μ i is the mean value of the feature x i σ i is the standard deviation of this feature.
[0088] The above-mentioned steps of data standardization processing can avoid the deviation of features with different dimensions on model training, and help to improve the stability and convergence speed of the model. Especially for the complex constraints and multi-dimensional features in the large-scale unit commitment problem, the standardization processing can effectively improve the solution accuracy.
[0089] In some embodiments, the label information is divided into a first difficulty label, a second difficulty label, and a third difficulty label, and the time nodes are divided into a first time threshold and a second time threshold. The steps of determining the label information matching the system clearing historical data according to the solution time of the corresponding solution of each system clearing historical data and the preset label annotation model include:
[0090] Judge whether the solution time of the corresponding solution of the system clearing historical data is less than the first time threshold;
[0091] If the solution time is less than the first time threshold, then match the system clearing historical data with the first difficulty label;
[0092] If the solution time is greater than or equal to the first time threshold, then judge whether the solution time is less than or equal to the second time threshold, and the second time threshold is greater than the first time threshold;
[0093] If the solution time of the said solution is greater than or equal to the first time threshold and less than or equal to the second time threshold, then match the system clearing historical data with the second difficulty label;
[0094] If the solution time of the said solution is greater than or equal to the second time threshold, then match the system clearing historical data with the third difficulty label.
[0095] In this embodiment, the first time threshold can be set to 500 seconds, and the second time threshold can be set to 1500 seconds. The first difficulty label can correspond to simple problems, the second difficulty label can correspond to medium difficulty, and the third label can correspond to difficult problems.
[0096] Suppose the system clearing historical data contains N clearing cases, and the feature vector of each clearing case is represented by xi. Taking the solution time as the target variable, the solution time can be represented by t i The sample data set can be defined as:
[0097] D = {(x 1 , t 1 ), (x 2 , t 2 ),..., (x N , t N )}.
[0098] In the formula, M is the dimension of the feature.
[0099] Among them, as Figure 2 shown, in the step of extracting features and labeling tags from the system clearing historical data to establish a sample data set with tag information, the features extracted from the system clearing historical data can include system data, power plant data, and unit data. The system data can include system load parameters and system reserve parameters. The power plant data can include power plant power parameters, power plant electricity parameters, power plant reserve parameters, etc. The unit data includes unit electricity parameters, unit power parameters, and unit operation parameters.
[0100] Then, generate tag information according to the solution time of each clearing case. Among them, the tag information can be represented by d i , and the value range is {0, 1, 2}, corresponding to the first difficulty label, the second difficulty label, and the third difficulty label respectively. The solution time can be represented by t i . The tag labeling model can include the following formula:
[0101]
[0102] That is to say, if the solution time of the said solution is less than the first time threshold (e.g., 500 seconds), then mark the system clearing historical data as the first difficulty label; if the solution time of the said solution is greater than or equal to the first time threshold (e.g., 500 seconds) and less than or equal to the second time threshold (e.g., 1500 seconds), then mark the system clearing historical data as the second difficulty label; if the solution time of the said solution is greater than or equal to the second time threshold (e.g., 1500 seconds), then mark the system clearing historical data as the third difficulty label. Of course, in other embodiments of the present invention, the first time threshold and the second time threshold can also be set to other values, which will not be elaborated here one by one.
[0103] In some embodiments, the steps of processing the sample data in the sample data set based on the extremely randomized forest algorithm to establish an extremely randomized forest prediction model include:
[0104] According to the label information of the sample data, divide the sample data in the sample data set into multiple subsets;
[0105] Based on the extremely randomized forest algorithm, process each of the subsets to establish an extremely randomized forest decision tree. There are at least two extremely randomized forest decision trees, and each extremely randomized forest decision tree corresponds to a subset;
[0106] Construct an extremely randomized forest prediction model according to all the extremely randomized forest decision trees.
[0107] That is to say, in step S30, first, according to the label information of the sample data, divide all the sample data into multiple smaller subsets, and ensure that each subset contains approximately or the same amount of label information; then use the extremely randomized forest algorithm to construct a corresponding extremely randomized forest decision tree for each subset; finally, integrate all the extremely randomized forest decision trees to form an extremely randomized forest prediction model, so that it can make full use of the prediction results of each extremely randomized forest decision tree, and finally obtain an accurate prediction result, thereby reducing the risk of overfitting and improving the stability and accuracy of the prediction at the same time.
[0108] In this embodiment, during the process of establishing the extremely randomized decision tree in step, it can be assumed that the feature set is X = {x 1 ,x 2 ,…,x M}, and the construction of each decision tree randomly extracts a subset from the training set and the corresponding label When constructing each tree node, the model randomly selects a subset from all the features Then select an optimal splitting feature f in this subset *, to minimize the impurity of the nodes. The objective function for node splitting can be expressed by the following formula:
[0109]
[0110] where T is the sample of the current node; T L is the sample of the left child node after splitting; T R is the sample of the right child node after splitting; G(T / T L / T R ) represents the impurity (such as Gini index or variance) of node T or T L or T R . Among them, the impurity G(T) of node T can be expressed by the following formula:
[0111]
[0112] where p k represents the proportion of samples belonging to the k-th class label in the current node.
[0113] Furthermore, after the step of constructing the extremely randomized forest prediction model, the power system unit scheduling optimization method further includes the steps of:
[0114] According to the multi-fold validation algorithm, divide the divided subsets into training subsets and validation subsets, and there are at least two training subsets;
[0115] According to the sample data in the sample data set, establish true positive samples, false positive samples, false negative samples and true negative samples;
[0116] Supplement the true positive samples, false positive samples, false negative samples and true negative samples to the validation subset to establish a validation test set;
[0117] Substitute the sample data in the validation test set into the extremely randomized forest prediction model, and calculate the model evaluation parameters. The model evaluation parameters include model accuracy, macro-average precision, macro-average recall, macro-average F1 score and macro-average area under the receiver operating characteristic curve. The macro-average F1 score is the macro-average value after the harmonic mean of precision and recall;
[0118] Use the model evaluation parameters and a preset overall reliability evaluation model to calculate the reliability score of the extremely randomized forest prediction model;
[0119] If the reliability score and the model evaluation parameters meet the preset conditions, it is considered that the extremely randomized forest prediction model runs reliably. The preset conditions include that at least one of the reliability score is greater than or equal to the score threshold, the macro-average recall rate is greater than or equal to the recall rate threshold, and the macro-average area under the receiver operating characteristic curve is greater than or equal to the area threshold under the receiver operating characteristic curve.
[0120] In this embodiment, during the process of verifying and testing the extremely randomized forest prediction model using the multi-fold cross-validation algorithm, the sample data set can be divided into 5 subsets, where 4 subsets are used for training and 1 subset is used for verification, so that all subsets can be tested in turn. The multi-fold cross-validation algorithm includes an objective function aimed at minimizing the prediction error of the model. The objective function can be expressed by the following formula:
[0121]
[0122] where K = 5 is the number of folds, is the predicted label of the model, d is the actual label, and T k is the validation set.
[0123] It should be noted that a true positive sample is a sample data with a difficulty level of k in the label information and actually being of this difficulty level. A false positive sample is a sample data with a difficulty level of k in the label information but actually not being of this difficulty level. A false negative sample is a sample data with an actual difficulty level of k but with a different difficulty level in the label information. A true negative sample is a sample data with an actual difficulty level not being k and the difficulty level in the label information not being k either.
[0124] In the above steps, these true positive samples, false positive samples, false negative samples, and true negative samples are supplemented to the validation subset, so that a confusion matrix can be constructed between the prediction result and the true result, and cross-comparison can be carried out by substituting them into the extremely randomized forest prediction model to ensure the accuracy of the evaluation.
[0125] In the specification of this application, to illustrate the specific process of calculating the model evaluation parameters and the reliability score, each parameter in the calculation process is described:
[0126]
[0127] The model accuracy rate is used to represent the proportion of correct predictions of the model and is a measure of the overall classification effect, which can measure the overall performance of the model. The model accuracy rate can be expressed by the following formula:
[0128]
[0129] The precision is used to represent the proportion of true positive samples in the model prediction. The macro-average precision is the precision after taking the macro-average for all classes. The precision can be expressed by the following formula:
[0130]
[0131] The recall is used to represent the proportion of samples that are actually of a certain class and are correctly classified. The macro-average recall is the recall after taking the macro-average for all classes. The recall can be expressed by the following formula:
[0132]
[0133] The F1 score is used to represent the harmonic mean of precision and recall. The macro-average F1 score is the F1 score after taking the macro-average for all classes. The F1 score can be expressed by the following formula:
[0134]
[0135] The area under the receiver operating characteristic curve (Area Under Curve, AUC) can measure the ability of the model to distinguish between samples of a certain class and samples of other classes. The macro-average area under the receiver operating characteristic curve is the area value under the receiver operating characteristic curve after taking the macro-average. It should be noted that the macro-average area under the receiver operating characteristic curve can also be called the macro-average AUC.
[0136] The reliability score can be expressed by the following formula:
[0137]
[0138] It should be noted that for large-scale unit commitment problems, α 1 = 0.2, α 2 = 0.2, α 3 = 0.2, α 4 = 0.2, α 5 = 0.2, that is, each weight is equal, to avoid an excessive influence of a certain parameter on the reliability score.
[0139] In this embodiment, the preset conditions include that the reliability score is greater than or equal to the score threshold, the macro-average recall is greater than or equal to the recall threshold, and the macro-average area under the receiver operating characteristic curve is greater than or equal to the area threshold under the receiver operating characteristic curve. That is, the reliability score, the macro-average recall, and the macro-average area under the receiver operating characteristic curve should all meet the preset conditions.
[0140] For ease of explanation, the scoring threshold can be set to 0.85, the recall rate threshold can be set to 0.8, and the area under the receiver operating characteristic curve threshold can be set to 0.85. When S model ≥0.85, F R (k)≥0.8, and F AUC ≥0.85 are all satisfied, it can be considered that the model has sufficient reliability and generalization ability. If some of the model's indicators are low (for example, the recall rate of difficult problems is lower than 70%), the model parameters can be adjusted (such as increasing the number of decision trees and adjusting the size of the feature subset), the sample distribution can be rebalanced (increasing the sample weight of difficult problems), or new features can be introduced to improve the classification performance.
[0141] It should be noted that in some other embodiments of the present invention, the scoring threshold, the recall rate threshold, and the area under the receiver operating characteristic curve threshold can all be set to other values, which will not be elaborated here.
[0142] In some embodiments, the step of dividing the sample data in the sample data set into multiple subsets according to the label information of the sample data includes:
[0143] Classify the sample data in the sample data set according to the label information of the sample data;
[0144] Based on the stratified sampling algorithm, divide the classified sample data into multiple subsets, and there are at least two subsets.
[0145] Specifically, first count the number of samples corresponding to each label information in the sample data set. The number of samples with the first difficulty label is N 0 ; the number of samples with the second difficulty label is N 1 ; the number of samples corresponding to the third difficulty label is N 2 , then the distribution ratio corresponding to each label information can be expressed by the following formula:
[0146]
[0147] In the process of dividing the classified sample data into multiple subsets using the stratified sampling algorithm in the step, assuming the proportion of the training subset is P train =0.7, and the proportion of the test subset is P test =0.3, then the number of samples in the training subset and the test subset are respectively: N train,k =p train ·N k , N test,k =p test ·N k .
[0148] Moreover, in the training subset, the number of samples corresponding to the first difficulty label, the second difficulty label, and the third difficulty label should each be maintained at N train,0 , N train,1 , N train,2 , so as to prevent the model from overfitting or underfitting a certain category during training.
[0149] In this embodiment, the unit scheduling solver includes an objective function and constraint conditions with the goal of minimizing the sum of start-stop costs and operating costs. The constraint conditions include at least one of a power demand constraint condition, a unit start-stop constraint condition, a unit power output constraint condition, a start-stop change constraint condition, a running ramp constraint condition, and a minimum running time constraint condition. It should be noted that during the operation of the unit scheduling solver, the corresponding constraint conditions can be selected according to the specific usage scenario.
[0150] In this embodiment, the objective function can be expressed by the following formula:
[0151]
[0152] In the formula, is the start-stop cost of unit i at time t for starting or shutting down. This cost depends on the start-stop state u of unit i at the current time t i,t and the start-stop state u at the previous time i,t - 1 . If the unit changes from shutdown to startup or from startup to shutdown, the corresponding startup or shutdown cost will be generated. is the operating cost of unit i at time t, which is usually proportional to the power generation p i,t of the unit. u i,t ∈{0, 1} is the start-stop decision of unit i at time t, where 1 represents startup and 0 represents shutdown. p i,t is the power generation of unit i at time t.
[0153] Among them, the power demand constraint condition means that at each time t, the total power generation of the system must meet the power demand D t . The power demand constraint condition can include the following formula:
[0154]
[0155] In the formula: D t is the power demand at time t. p i,t is the power generation of unit i at time t.
[0156] The unit start-stop constraint condition means that each unit can only be in the "on" or "off" state in the start-stop decision. The unit start-stop constraint condition may include the following formula:
[0157]
[0158] The unit power output constraint condition means that the generated power p of each unit i,t must be between its minimum generated power and maximum generated power. The unit power output constraint condition may include the following formula:
[0159]
[0160] In the formula: and respectively represent the minimum and maximum generated powers of unit i at time t.
[0161] The unit operation ramp constraint condition means that the rate of power change cannot exceed a certain threshold. The unit operation ramp constraint condition may include the following formula:
[0162]
[0163] In the formula, p i,t is the generated power of unit i at time t; is the generated power of unit i at time t-1. Δp i is the maximum allowable power change of unit i between time t and time t-1.
[0164] The minimum operation time constraint condition of the generator set is used to limit that the unit needs to run continuously for at least a certain duration after starting up before it can be shut down. The minimum operation time constraint condition of the generator set may include the following formula:
[0165]
[0166] In the formula, m represents the minimum operation time.
[0167] In some embodiments, the solver parameters at least include the maximum number of iterations, convergence accuracy, allowable relative error, number of parallel processing threads, heuristic algorithm switch, and setting of the initial relaxation variable. Among them, the correspondence between the label information and the solution parameters can be represented by the following table:
[0168]
[0169]
[0170] In addition, it should be clear that Figure 1The flow schematic diagram of the power system unit scheduling optimization method shown therein. The order of steps is only shown based on the logical order provided by one of the numerous embodiments of the present invention. This order aims to clearly illustrate the execution process of the method in a specific application scenario. However, in practical applications, the flexibility and scalability of the present invention allow technicians to adjust the order of the above steps according to different actual requirements and environmental conditions. For example, there are parallel or juxtaposed relationships between some steps, which will not be elaborated one by one here.
[0171] The power system unit scheduling optimization method provided in an embodiment of the present invention is applicable to a variety of scenarios. In the following, three scenarios will be described in the specification of this application:
[0172] Scenario 1: The scenario of the unit commitment problem in a medium-sized power system. Assume that the system scale is about 158 generating units, mainly composed of coal-fired units and gas-fired units, and a small amount of wind power and photovoltaic power are connected to the grid. The daily load curve of the system changes smoothly, the peak-valley difference is about 20%, and the requirement for reserve capacity configuration is moderate. After being processed by the power system unit scheduling optimization method provided by the present invention in this scenario, the maximum number of iterations is set to 600, the convergence accuracy is 1e-4, and the number of threads is 1 to 4.
[0173] Scenario 2: The scenario of the unit commitment problem in a large power system. Assume that the system scale includes 835 generating units, and the proportion of new energy units reaches 14.8%. After being processed by the power system unit scheduling optimization method provided by the present invention in this scenario, the maximum number of iterations is 2000, the convergence accuracy is 1e-5, and the number of threads is set to 4 to 8 to improve the calculation efficiency and accuracy.
[0174] Scenario 3: The scenario of the unit commitment problem in an extra-large cross-regional power system. Assume that the system scale includes 2181 generating units, and the new energy penetration ratio is close to 29.7%. After being processed by the power system unit scheduling optimization method provided by the present invention in this scenario, the maximum number of iterations is set to 10000, the convergence accuracy is 1e-6, and the number of threads is 8 to 16 to ensure efficient solution under large-scale data and complex constraint conditions.
[0175] The conclusions after using the power system unit scheduling optimization method provided by the present invention in the above three scenarios are as follows:
[0176] In Scenario 1, the scheduling cost using the power system unit scheduling optimization method of the present invention is 35.958 million yuan, compared with 35.961 million yuan of the traditional method, the error is 0.0084%, and the solution time is reduced from 1056.5 seconds of the traditional method to 950.4 seconds, a reduction of about 10.0%.
[0177] In Scenario 2, the scheduling cost of the power system unit scheduling optimization method of the present invention is 151.195 million yuan. Compared with the 151.213 million yuan of the traditional method, the error is 0.01%. The solution time is reduced from 3325.1 seconds of the traditional method to 2765.4 seconds, a reduction of approximately 16.8%.
[0178] In Scenario 3, the scheduling cost of the power system unit scheduling optimization method of the present invention is 649.918 million yuan. Compared with the 650.117 million yuan of the traditional method, the error is 0.03%. The solution time is reduced from 6154.2 seconds of the traditional method to 4561.9 seconds, a reduction of approximately 25.9%.
[0179] Generally speaking, the power system unit scheduling optimization method of the present invention is highly accurate compared with the traditional method and significantly reduces the solution time. Specifically, compared with the traditional method, the average reduction in the solution time of the power system unit scheduling optimization method is approximately 17.6%, verifying the computational efficiency advantage of the new method in large-scale unit commitment problems.
[0180] Reference Figure 3 , is a schematic structural diagram of a power system unit scheduling optimization device 100 provided by an embodiment of the present invention. The power system unit scheduling optimization device 100 includes an acquisition module 110, a sample annotation module 120, a model construction module 130, a label matching module 140, a parameter matching module 150, and a solution determination module 160. The above modules are electrically connected to each other to achieve information transmission and reception. Of course, in some embodiments, the acquisition module 110, the sample annotation module 120, the model construction module 130, the label matching module 140, the parameter matching module 150, and the solution determination module 160 may also be integrated into one body to make the overall structure more compact.
[0181] Among them, the obtaining module 110 is used to obtain the system clearing historical data and the clearing data to be processed. The sample annotation module 120 is used to extract features and label the system clearing historical data, and establish a sample data set with label information. The model construction module 130 is used to process the sample data in the sample data set according to the extremely randomized forest algorithm, and establish an extremely randomized forest prediction model. The label matching module 140 is used to substitute the clearing data to be processed into the extremely randomized forest prediction model to obtain the label information matching the clearing data. The parameter matching module 150 is provided with a parameter matching model, and the parameter matching module 150 is used to determine the solver parameters matching the clearing data according to the label information of the clearing data. The solution determination module 160 is provided with a unit scheduling solver, and the solution determination module 160 is used to process the clearing data and the solver parameters by using the unit scheduling solver to obtain a unit scheduling solution matching the clearing data.
[0182] It should be noted that the power system unit scheduling optimization device provided in the embodiment of the present application has the same implementation principle and the same technical effects as those in the foregoing embodiment of the power system unit scheduling optimization method. For the sake of brief description, for the parts not mentioned in the device embodiment, reference may be made to the corresponding content in the foregoing embodiment of the power system unit scheduling optimization method.
[0183] Reference Figure 4 , which is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. The electronic device provided in this embodiment includes a processor 210 and a memory 220. The memory 220 stores machine-readable instructions executable by the processor 210. When the machine-readable instructions are executed by the processor 210, the steps in the power system unit scheduling optimization method described in any of the foregoing embodiments are executed. Among them, there is at least one of the processor 210 and the memory 220.
[0184] In this embodiment, the electronic device further includes a communication interface 230 and a communication bus 240. Among them, the processor 210, the memory 220, and the communication interface 230 are connected to each other through the communication bus 240. The communication bus 240 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 3 only a thick line is used to represent the communication bus 240 in, but it does not mean that there is only one communication bus 240 or one type of communication bus 240. The processor 210 can also be called a controller, and there is no limitation on the name.
[0185] In the embodiments of the present application, the memory 220 stores instructions that can be executed by at least one processor 210. By executing the instructions stored in the memory 220, the at least one processor 210 can execute the steps in the power system unit scheduling optimization method described above. The processor 210 can implement Figure 3 the functions of each module in the device shown.
[0186] Among them, the processor 210 is the control center of the device. It can connect various parts of the entire control device through various interfaces and lines. By running or executing the instructions stored in the memory 220 and calling the data stored in the memory 220, the various functions of the device and process data, so as to monitor the device as a whole.
[0187] In a possible design, the processor 210 may include one or more processing units. The processor 210 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, the operating body interface, and application programs, etc. The modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor either. In some embodiments, the processor 210 and the memory 220 may be implemented on the same chip or separately on independent chips.
[0188] The processor 210 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the power system unit scheduling optimization method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware processor or completed by a combination of hardware and software modules in the processor.
[0189] The memory 220, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 220 may include at least one type of storage medium, for example, it may include flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disk, etc. The memory 220 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 220 in the embodiments of the present application may also be a circuit or any other device that can implement the storage function, for storing program instructions and / or data.
[0190] By programming the design of the processor 210, the code corresponding to the power system unit scheduling optimization method introduced in the foregoing embodiments can be solidified into the chip, so that the chip can execute Figure 2 the steps of the power system unit scheduling optimization method of the illustrated embodiment. How to program the design of the processor 210 is a well-known technology to those skilled in the art and will not be elaborated here.
[0191] The embodiments of the present application also provide a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by the processor 210, they are used to implement the power system unit scheduling optimization method described in any of the foregoing embodiments. Therefore, it will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer storage medium involved in the present invention, please refer to the description of the method embodiments of the present invention.
[0192] In some possible implementation manners, each aspect of the power system unit scheduling optimization method provided in the present application may also be implemented in the form of a program product, which includes program code. When the program product runs on the device, the program code is used to cause the control device to execute the steps in the power system unit scheduling optimization method according to various exemplary implementation manners of the present application described above in this specification.
[0193] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0194] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0195] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0196] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0197] In addition, any process or method description in the flowchart or described in other ways herein can be understood to represent a module, segment, or part of code including one or more executable instructions for implementing a customized logic function or process. The scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed. This should be understood by those skilled in the art to which the embodiments of the present invention belong.
[0198] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for optimizing the dispatching of power system units, characterized in that: The steps include: Obtain system clearing historical data and pending clearing data; Extract features and label the historical data of the system clearing to establish a sample data set with label information; Based on the extreme random forest algorithm, the sample data in the sample data set is processed to establish an extreme random forest prediction model; Substituting the clearing data to be processed into the extreme random forest prediction model to obtain label information matching the clearing data; Determining solver parameters that match the clearing data based on the tag information and a preset parameter matching model; Substitute the clearing data and the solver parameters into a preset unit scheduling solver to obtain a unit scheduling solution that matches the clearing data.
2. The method for optimizing dispatching of power system units according to claim 1, characterized in that: The step of processing the sample data in the sample data set based on the extreme random forest algorithm to establish an extreme random forest prediction model includes: Dividing the sample data in the sample data set into a plurality of subsets according to the label information of the sample data; Based on the extreme random forest algorithm, each of the subsets is processed to establish an extreme random forest decision tree, wherein there are at least two extreme random forest decision trees, and each extreme random forest decision tree corresponds to one subset; Based on all the extreme random forest decision trees, an extreme random forest prediction model is constructed.
3. The method for optimizing dispatching of power system units according to claim 2, characterized in that: The power system unit dispatch optimization method further comprises the following steps after the step of constructing the extreme random forest prediction model: According to a multi-fold validation algorithm, the divided subsets are divided into a training subset and a validation subset, and there are at least two training subsets; According to the sample data in the sample data set, establishing true positive samples, false positive samples, false negative samples and true negative samples; Adding the true positive samples, false positive samples, false negative samples and true negative samples to the validation subset to establish a validation test set; Substituting the sample data in the validation test set into the extreme random forest prediction model, and calculating the model evaluation parameters, the model evaluation parameters include model accuracy, macro-average precision, macro-average recall, macro-average F1 score and macro-average area under the receiver operating characteristic curve, the macro-average F1 score is the macro-average value after the harmonic average of the precision and the recall; The model evaluation parameters and the preset overall reliability evaluation model are used to calculate the reliability score of the extreme random forest prediction model; If the reliability score and the model evaluation parameters meet the preset conditions, the extreme random forest prediction model is considered to be reliable, and the preset conditions include at least one of the reliability score being greater than or equal to the score threshold, the macro-average recall rate being greater than or equal to the recall rate threshold, and the macro-average area under the receiver operating characteristic curve being greater than or equal to the area threshold under the receiver operating characteristic curve.
4. The method for optimizing dispatching of power system units according to claim 2, characterized in that: The step of dividing the sample data in the sample data set into a plurality of subsets according to the label information of the sample data comprises: Classifying the sample data in the sample data set according to the label information of the sample data; Based on a stratified sampling algorithm, the classified sample data is divided into a plurality of subsets, and there are at least two subsets.
5. The method for optimizing dispatching of power system units according to claim 1, characterized in that: The system clearing historical data includes operation data and corresponding solution time. The steps of extracting features and labeling the system clearing historical data to establish a sample data set with label information include: Extracting features from the operating data in the historical data cleared by the system, and preprocessing the extracted features, wherein the preprocessing at least includes filling missing values, removing outliers, and data standardization; According to the solution solution time corresponding to each system clearing historical data and a preset labeling model, determine the label information matching the system clearing historical data, wherein the labeling model includes a plurality of label information and a plurality of time nodes, the plurality of time nodes gradually increase, and the label information corresponds to the time nodes one by one; A sample data set with label information is established according to the operation data, the solution solution time and the label information.
6. The method for optimizing dispatching of power system units according to claim 5, characterized in that: The label information is divided into a first difficulty label, a second difficulty label and a third difficulty label, and the time node is divided into a first time threshold and a second time threshold; the step of determining the label information matching the system clearing historical data according to the solution solution time corresponding to each system clearing historical data and the preset label annotation model comprises: Determine whether the solution time corresponding to the historical data cleared by the system is less than a first time threshold; If the solution time is less than the first time threshold, matching the system clearing historical data with the first difficulty label; If the solution time is greater than or equal to the first time threshold, determining whether the solution time is less than or equal to the second time threshold, and the second time threshold is greater than the first time threshold; If the solution time of the solution is greater than or equal to the first time threshold and less than or equal to the second time threshold, matching the system clearing historical data with the second difficulty label; If the solution time is greater than or equal to the second time threshold, the system clearing historical data is matched with a third difficulty tag.
7. The method for optimizing dispatching of power system units according to claim 1, characterized in that: The unit scheduling solver includes an objective function and constraints with the goal of minimizing the sum of the start-stop cost and the operating cost, wherein the constraints include at least one of a power demand constraint, a unit start-stop constraint, a unit power output constraint, a start-stop change constraint, an operating ramp constraint, and a minimum operating time constraint; And / or, the solver parameters at least include settings of a maximum number of iterations, convergence accuracy, allowable relative error, number of parallel processing threads, a heuristic algorithm switch, and initial slack variables.
8. A power system unit dispatch optimization device, characterized in that: include: An acquisition module, the acquisition module is used to acquire system clearing historical data and clearing data to be processed; A sample labeling module, which is used to extract features and label the system clearing historical data, and establish a sample data set with label information; A model building module, wherein the model building module is used to process the sample data in the sample data set according to the extreme random forest algorithm to establish an extreme random forest prediction model; A label matching module, wherein the label matching module is used to substitute the clearing data to be processed into the extreme random forest prediction model to obtain label information matching the clearing data; A parameter matching module, wherein a parameter matching model is provided in the parameter matching module, and the parameter matching module is used to determine solver parameters matching the clearing data according to the label information of the clearing data; A plan establishment module is provided with a unit scheduling solver, and the plan establishment module is used to use the unit scheduling solver to process the clearing data and the solver parameters to obtain a unit scheduling plan that matches the clearing data.
9. An electronic device, characterized in that: It comprises a processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and when the machine-readable instructions are executed by the processor, the steps in the power system unit dispatch optimization method described in any one of claims 1-7 are executed.
10. A computer-readable storage medium, characterized in that: Machine-readable instructions are stored therein. When the storage medium is run on an electronic device, the machine-readable instructions are used to enable the electronic device to execute the steps in the power system unit dispatch optimization method described in any one of claims 1-7.