A power prediction method, system, device and storage medium based on sample clustering
Through the power prediction method based on sample clustering, the DTW distance calculation formula is improved, isolated or weak correlation sample points are eliminated, and the power prediction model is established, which solves the problem of insufficient prediction accuracy of wind power and photovoltaic power in the existing technology, and achieves higher prediction accuracy and lower operating costs.
Patent Information
- Application Number
- CN202510424282.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The existing technology is difficult to achieve high-accurate wind power and photovoltaic power prediction, resulting in unreasonable new energy consumption space and affecting the safe and stable operation of the power system.
Using a power prediction method based on sample clustering, by obtaining the time series data of the historical samples of the power system, preprocessing and feature construction, calculating the correlation between the preferred features and the predicted target value, improving the DTW distance calculation formula, performing sample point clustering, eliminating isolated or weak correlation sample points, and establishing a power prediction model.
It improves the accuracy of power prediction, reduces computing power requirements, saves operating costs, enhances the prediction capabilities of many physical quantities time series in the power industry, and supports the improvement of power balance in the power grid.
Smart Images

Figure CN119944674B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to power system prediction, and more particularly to a power prediction method, device, equipment and storage medium based on sample clustering. Background Art
[0002] In recent years, new energy represented by photovoltaic and wind power will gradually replace traditional energy and show good development prospects. Due to the randomness and volatility of wind power and photovoltaic power, their large-scale access to the power grid has brought impacts on the safe and stable operation of the power system.
[0003] Due to the characteristics that electric energy is difficult to store in large quantities and the power demand changes at all times, etc., it is required that the system power generation should achieve dynamic balance with the change of load. Predicting the future output power of wind power and photovoltaic power in advance and reserving accommodation space according to the prediction results are important technical means to improve the new energy accommodation level and ensure the safety of the power system. The power prediction accuracy directly affects whether the reserved accommodation space is reasonable. In the case of poor prediction accuracy, in order to ensure the system safety, the standby of conventional power sources increases, squeezing the new energy accommodation space, thus causing an increase in the amount of abandoned wind and light. Establishing a high-accuracy wind power and photovoltaic power prediction system and scientifically formulating a dispatching plan and reserving a reasonable accommodation space according to the prediction results are effective measures to improve the new energy accommodation capacity.
[0004] At the same time, as the basis of safe and stable dispatching, the accuracy and stability of load prediction will further improve the automation level of the power grid dispatching system and the degree of economic dispatching; it will become an important basis for formulating the clearing plan of the power spot market system and affect the reliability of the clearing plan; it will become the solid foundation of the new generation of power grid intelligent dispatching system and provide reliable up-to-date load demand information for power grid decision-making.
[0005] In addition, the smooth progress of power market bidding transactions must be based on a scientific and reasonable bidding algorithm. As the core issue of power market transactions, the ultimate goal of the bidding algorithm is to improve its economy under the condition of safe and stable operation of the system. In this process, the accurate prediction of the clearing price plays an important role.
[0006] In addition to the above prediction requirements, the prediction of many other physical quantities in the power industry also has practical significance for the safe and stable operation of the power grid. All these physical quantities have the characteristics of time series and periodicity. In this case, the accurate prediction and analysis of time series become an urgent problem to be solved. Summary of the Invention
[0007] Object of the Invention: Aiming at the above-mentioned drawbacks, the present invention provides a power prediction method, device, equipment and storage medium based on sample clustering to improve prediction accuracy.
[0008] Technical solution: To solve the above problems, the present invention adopts a power prediction method based on sample clustering, including the following steps:
[0009] Obtain the time series data set of the historical samples of the power system, preprocess the time series data set, and extract the basic features. Select the important features from the basic features, and construct derivative features from the important features. Use the important features and derivative features as seed features, randomly select the remaining features in the basic features except the important features and put them into several sub-feature sets, and perform preliminary feature selection on each sub-feature set respectively to obtain preliminary features; perform feature selection on the preliminary features and seed features to obtain the preferred features;
[0010] Calculate the correlation between each preferred feature in the historical samples of the power system and the predicted target value, and use the correlation and the absolute value of the first-order differential of the feature value of the preferred feature as weights to improve the DTW distance calculation formula between sample points;
[0011] According to the improved DTW distance calculation formula, cluster the sample points, eliminate the isolated or weakly correlated sample points, use the data set after eliminating the isolated or weakly correlated sample points as the training set and the test set, train the power prediction model through the training set and the test set, and perform power prediction through the power prediction model.
[0012] Further, the preprocessing of the time series data set includes introducing a sliding window, calculating the standard deviation of a single feature within the sliding window, and eliminating the sample points with a standard deviation less than The window width of the sliding window is the number of consecutive constant values.
[0013] Further, the improved DTW distance calculation formula is:
[0014] ;
[0015] ;
[0016] Among them, is the i-th sample point in time series A, is the j-th sample point in time series B, , is a specified constant, is the feature value of the preferred feature in the sample point is the feature value of the preferred feature in the sample point is the preferred feature is the feature value of is the preferred feature is the correlation between and the predicted target value, is the sample point is the preferred feature in the sample point The first-order differential of the eigenvalue and the sample points Preferred features among them The absolute value of the difference between the first-order differentials of the eigenvalues, is the total number of preferred features.
[0017] Furthermore, the measurement criteria for the correlation between the preferred features and the predicted target value include using the maximum information coefficient MIC.
[0018] Furthermore, when clustering the sample points, the weighted mean of all sample points in each cluster in the clustering result is calculated respectively, and the weighted mean of all sample points in the data set is calculated. The sample points in the clusters whose weighted mean exceeds the threshold range are identified as isolated or weakly correlated sample points and are excluded;
[0019] The calculation formula for the weighted mean of the cluster is:
[0020] ;
[0021] where, is the number of sample points in the cluster in the cluster is the cluster the preferred feature of the j-th sample point in the cluster is the eigenvalue of the preferred feature is the preferred feature of the j-th sample point is the first-order differential of the eigenvalue of the preferred feature.
[0022] Furthermore, the preliminary feature selection includes filter selection and wrapper selection. When using filter selection, the feature correlation measurement criteria adopted include the combination of the maximum information coefficient MIC and the Pearson correlation coefficient, and the features with the highest maximum information coefficient MIC and Pearson correlation coefficient are alternately selected.
[0023] The present invention also adopts a power prediction system based on sample clustering, including:
[0024] A feature selection module, used to obtain the historical sample data set of the power system, preprocess the data set, extract basic features, select important features from the basic features, construct derivative features from the important features, use the important features and derivative features as seed features, randomly extract the remaining features in the basic features except the important features into several sub-feature sets, perform preliminary feature selection on each sub-feature set respectively to obtain preliminary features, and perform feature selection on the sub-preliminary features and seed features to obtain preferred features;
[0025] A sample clustering module, which is used to calculate the correlation between each preferred feature in the historical samples of the power system and the predicted target value, and use the correlation and the absolute value of the first-order differential of the preferred feature value as weights to improve the DTW distance calculation formula between sample points; according to the improved DTW distance calculation formula, cluster the sample points, and eliminate isolated or weakly correlated sample points;
[0026] A prediction module, which is used to use the data set after eliminating isolated or weakly correlated sample points as the training set and the test set, train the power prediction model through the training set and the test set, and perform power prediction through the power prediction model.
[0027] The present invention also adopts a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are realized.
[0028] The present invention also adopts a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are realized.
[0029] Beneficial effects: Compared with the prior art, the significant advantage of the present invention is that it adopts a feature grouping method to save operating costs, and a sample clustering algorithm to improve the prediction accuracy of time series in the power industry for noise reduction processing. The clustering distance calculation reflects the importance and different change trends of different preferred features by introducing correlation and change rate, improves the importance of important features in the prediction weight, and thus improves the prediction accuracy. Reducing the computing power requirement through feature grouping saves operating costs. The noise reduction processing based on sample clustering predicts the time series of many physical quantities in the power industry, improves the prediction accuracy, and can be used as a powerful supplement and alternative to the existing prediction and analysis methods. The present invention improves the prediction ability of the time series of many physical quantities in the power industry, and provides important technical support for promoting the improvement of the power grid power balance in the power industry. Description of the Drawings
[0030] Figure 1 It is a schematic flow chart of the prediction method in the present invention.
[0031] Figure 2 It is a principle block diagram of feature grouping and feature selection in the present invention. Detailed Embodiments
[0032] Embodiment 1
[0033] As Figure 1 shown, a power prediction method based on sample clustering in this embodiment. This prediction method mainly realizes clustering and purification of historical data and influence factor data. The implementation process is as follows:
[0034] Due to the possible diversity of data sources and technical and resource limitations during the data collection process, it is difficult to guarantee data quality. There are a large number of outlier and weakly correlated sample data in the dataset, which will lead to errors in the finally generated prediction model and a decrease in prediction accuracy. To address these issues, a dataset of historical samples is collected in advance, and various conventional preprocessings are performed on the data samples. Assume that the processed dataset obtained at this time is . On this basis, some time and physical features are selected as important features, and feature construction, namely feature engineering, is carried out based on these features. The features obtained are called derived features. There are many construction methods. Since it is a time series, first, the number of the natural day of each date in the current year, the number of the natural day of each month, and the number of the hour of each hour in the current day can be constructed. Secondly, select the features that are particularly important in terms of physical logic. The derived features can be the m-th power (m is an integer not equal to 0) of the important feature values, and the n-th order difference can also be constructed, as well as the time shift of the feature by n time points (n is a natural number). For new energy or load-related predictions, wind speed, wind direction, and irradiance are important features with strong physical logic correlations. Specifically, specify wind speed at 10m (ws10), wind speed at 30m (ws30), wind speed at 50 (ws50), wind speed at 70m (ws70), wind speed at 80m (ws80), wind speed at 100m (ws100), wind speed at 120m (ws120), specify wind direction at 10m (wd10), wind direction at 30m (wd30), wind direction at 50m (wd50), wind direction at 70m (wd70), wind direction at 80m (wd80), wind direction at 100m (wd100), wind direction at 120m (wd120), direct radiation (DirectR), diffuse radiation (DiffuseR). If it is a wind power-related prediction, the height range of wind speed and wind direction is usually from 10m to slightly higher than the height of the wind turbine (in this embodiment, there are 7 related to wind speed, 7 related to wind direction, and 2 related to irradiance, totaling 16). In addition, there is still a large amount of other meteorological data. For various physical features representing wind speed, wind speed, wind direction, and irradiance, the second power, third power, first-order difference, and second-order difference are constructed respectively, totaling 16 * 4 = 64.
[0035] Since there are a large number of all features including the above important features and derived features, feature selection is indispensable, which necessarily requires a large amount of computing power support. However, in current application sites, especially in the engineering sites of distributed new energy, this kind of computing power is usually lacking. Such as Figure 2As shown, here, the concepts of seed features and feature grouping are introduced. The above-mentioned important features and derivative features are designated as seed features, which may be denoted as SF (seed feature, and there are 64 SFs in this embodiment). For other features, a series of sub-feature sets are established respectively, with a total of M. Features are randomly selected as much as possible from all the remaining available features and placed into the M sub-feature sets respectively, and the number of features contained in each sub-feature is made as equal as possible.
[0036] Perform preliminary feature selection on each sub-feature set. Each sub-feature set obtains n preliminary features respectively, with a total of n*M preliminary features (in this embodiment, M = 3 and n = 4, so n*M = 12). The preliminary feature selection can use filter selection, wrapper selection or other methods, which will not be elaborated here. In this embodiment, the filter selection method is adopted, and the Maximum Information Coefficient (MIC) and Pearson Correlation Coefficient (PCC) are selected as the criteria for measuring the importance degree of feature correlation, that is: the first selection is the maximum information coefficient MIC, the second selection is the Pearson correlation coefficient PCC, and the feature with the highest score is selected alternately. If a situation where it has been selected appears, then the feature with the second highest score is selected (other correlation measurement criteria can be selected, and the number of correlation measurement criteria can also be multiple, as long as rotation is satisfied). After selecting M*n features from all feature subsets, if the number is still very large, such as greater than a set value L (L = 40 is taken in this embodiment), it can be considered to group again and repeat the above steps. Feature selection is performed on the selected n*M features and SF seed features to select the final preferred features (in this embodiment, M = 3, n = 4, SF = 64), and the number of preferred features can be determined according to the actual situation.
[0037] For the data set required for training, due to equipment failures, calculation errors or other reasons, there will be continuous constant values or approximately constant values due to superimposed noise (except when the new energy is fully dispatched or the photovoltaic output at night is 0). To eliminate these constant values, a sliding window is introduced, and the width of the window is the number of consecutive constant values. By calculating the standard deviation of a single feature within the sliding window, the sample points with a standard deviation close to 0 (which can be set to , a natural number) are eliminated.
[0038] In addition, there must be a large number of isolated points or data with weak correlation with the prediction target in the dataset. Cluster the sample points and remove the isolated or weakly correlated sample points. Since the DTW distance can better track the time offset and trend change differences, the DTW distance is selected as the measurement standard for cluster analysis. Before performing cluster analysis using DTW as the clustering distance, optimize the clustering distance.
[0039] In the calculation process of the DTW distance function, the calculation formula is as follows:
[0040] ;
[0041] where, is the i-th sample point in time series A, is the j-th sample point in time series B.
[0042] In the above formula, the elements of the distance matrix are measured using the Euclidean distance. To reflect the importance of different preferred features and the differences in trend changes, by introducing correlation and change rate, the calculation formula for measuring the distance between the elements of the above distance matrix is improved as follows:
[0043] ;
[0044] where, , is an arbitrarily specified constant used to proportionally adjust the value when the value is too small to affect the calculation accuracy during calculation, is the preferred feature of the sample point in the eigenvalue, is the preferred feature of the sample point in the eigenvalue, is the correlation between the preferred feature and the prediction target value, is the absolute value of the difference between the first-order differential of the eigenvalue of the preferred feature in the sample point and the first-order differential of the eigenvalue of the preferred feature in the sample point , is the total number of preferred features.
[0045] For any preferred feature in the dataset, calculate its correlation with the prediction target, i.e., the label, respectively. In this embodiment, MIC is used as the correlation measurement standard (other measurement standards can be selected according to different scenarios and effects), and a normalization operation is performed to obtain ; At the same time, select the first-order differential of the sample point on each preferred feature , as a change rate metric; thus, for any two sample points, calculate the absolute value of the difference in the first-order derivatives of each preferred feature of the two sample points respectively, and perform normalization to obtain .
[0046] Since the fitting result of the k-means model is often a circular cluster and the fitting effect for clusters of many other specific shapes is not ideal, the Gaussian mixture model (GMM) that can fit data distributions of arbitrary shapes is selected. Calculate the mean value weighted by the correlation metric and the first-order derivative for each cluster in the clustering result for all samples according to the preferred features and the prediction target, i.e., the label, that is: For all samples, according to the correlation metric between the preferred features and the prediction target, i.e., the label, and the mean value weighted by the first-order derivative, that is:
[0047] ;
[0048] where is the number of samples in cluster , is the eigenvalue of the preferred feature of the j-th sample point in cluster , is the first-order derivative of the eigenvalue of the preferred feature of the j-th sample point. Calculate the weighted mean value of all samples in the data set in the same way. If the mean value of a cluster far exceeds the mean value of the data set, then the data of this cluster is considered as an outlier or a sample with weak correlation. The exceeded range can be adjusted according to the differences of the data set. Establish a time series prediction model with the data set obtained after removing the outlier or weakly correlated sample data as the training set and the test set, and finally achieve the goal of significantly improving the prediction accuracy.
[0049] Example 2
[0050] A power prediction system based on sample clustering in this example includes:
[0051] A feature selection module, used to obtain the historical sample data set of the power system, preprocess the data set, extract basic features, select important features from the basic features, construct derivative features from the important features, use the important features and the derivative features as seed features, randomly extract the remaining features in the basic features except the important features into several sub-feature sets, perform preliminary feature selection on each sub-feature set respectively to obtain preliminary features, and perform feature selection on the sub-preliminary features and the seed features to obtain preferred features;
[0052] A sample clustering module, which is used to calculate the correlation between each preferred feature in the historical samples of the power system and the predicted target value, take the correlation and the absolute value of the first-order differential of the preferred feature value as weights, and improve the DTW distance calculation formula between sample points; according to the improved DTW distance calculation formula, cluster the sample points and eliminate isolated or weakly correlated sample points;
[0053] A prediction module, which is used to use the data set after eliminating isolated or weakly correlated sample points as the training set and the test set, train the power prediction model through the training set and the test set, and perform power prediction through the power prediction model.
[0054] Embodiment 3
[0055] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0056] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0057] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0058] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the steps specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps of the functions specified in one block or a plurality of blocks.
[0059] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope of the present invention as protected by the claims. These all fall within the protection scope of the present invention.
Claims
1. A power forecasting method based on sample clustering, characterized in that: The following steps are involved: Obtain a time series data set of historical samples of the power system, preprocess the time series data set, extract basic features, select important features from the basic features, perform feature construction on the important features to obtain derived features, use the important features and derived features as seed features, randomly extract the remaining features from the basic features except the important features and put them into several sub-feature sets, perform preliminary feature selection on each sub-feature set to obtain preliminary features; perform feature selection on the preliminary features and seed features to obtain preferred features; The correlation between each preferred feature and the predicted target value in the power system history samples is calculated, and the absolute value of the first-order differential of the correlation and the preferred feature value is used as the weight to improve the DTW distance calculation formula between sample points; The improved DTW distance calculation formula is: ; ; in, is the i-th sample point in time series A, is the jth sample point in time series B, , To specify a constant, For sample points Selected features The characteristic value of For sample points Selected features The characteristic value of The preferred feature The correlation between the predicted target value and For sample points Selected features First-order differential of eigenvalue and sample points Selected features The absolute value of the difference of the first differentials of the eigenvalues, is the total number of selected features; According to the improved DTW distance calculation formula, the sample points are clustered, and isolated or weakly correlated sample points are eliminated. The data set without isolated or weakly correlated sample points is used as the training set and the test set. The power prediction model is trained by the training set and the test set, and the power prediction model is used to perform power prediction.
2. The power forecasting method according to claim 1, characterized in that: The preprocessing of the time series data set includes: introducing a sliding window, calculating the standard deviation of a single feature in the sliding window, and removing the feature with a standard deviation less than The sample points of is a natural number, and the window width of the sliding window is the number of continuous constant values.
3. The power forecasting method according to claim 2, characterized in that: The metric of the correlation between the preferred features and the predicted target value includes using the maximum information coefficient MIC.
4. The power forecasting method according to claim 2, characterized in that: When clustering the sample points, the weighted mean of all sample points in each cluster in the clustering result is calculated respectively, and the weighted mean of all sample points in the data set is calculated, and the weighted mean of each cluster is compared with the weighted mean of the data set. The sample points in the clusters whose weighted means exceed the threshold range are identified as isolated or weakly correlated sample points and are removed; The calculation formula of the weighted mean of the cluster is: ; in, Cluster The number of sample points in , Cluster The optimal feature of the jth sample point The characteristic value of Select features for the jth sample point First-order differential of the eigenvalue.
5. The power forecasting method according to claim 1, characterized in that: The preliminary feature selection includes filtering selection and wrapping selection. When filtering selection is adopted, the feature correlation metric adopted includes a combination of maximum information coefficient MIC and Pearson correlation coefficient, and the features with the highest maximum information coefficient MIC and Pearson correlation coefficient are alternately selected.
6. A power prediction system based on sample clustering, characterized in that: include: The feature selection module is used to obtain a historical sample data set of the power system, preprocess the data set, extract basic features, select important features from the basic features, perform feature construction on the important features to obtain derived features, use the important features and the derived features as seed features, randomly extract the remaining features except the important features from the basic features and put them into several sub-feature sets, perform preliminary feature selection on each sub-feature set to obtain preliminary features, perform feature selection on the sub-preliminary features and the seed features to obtain preferred features; The sample clustering module is used to calculate the correlation between each preferred feature and the predicted target value in the historical samples of the power system, and use the absolute value of the first-order differential of the correlation and the preferred feature value as weights to improve the DTW distance calculation formula between sample points; according to the improved DTW distance calculation formula, the sample points are clustered to eliminate isolated or weakly correlated sample points; The improved DTW distance calculation formula is: ; ; in, is the i-th sample point in time series A, is the jth sample point in time series B, , To specify a constant, For sample points Selected features The characteristic value of For sample points Selected features The characteristic value of The preferred feature The correlation between the predicted target value and For sample points Selected features First-order differential of eigenvalue and sample points Selected features The absolute value of the difference of the first differentials of the eigenvalues, is the total number of selected features; The prediction module is used to use the data set with isolated or weakly correlated sample points removed as the training set and the test set, train the power prediction model through the training set and the test set, and perform power prediction through the power prediction model.
7. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Power grid line loss prediction method and device, storage medium and equipment
CN114781693A
Two-stage short-term power load prediction method based on feature selection and dimensionality reduction clustering
CN117713037A