A regional photovoltaic output scene division method, system, device and storage medium
Patent Information
- Application Number
- CN202310422767.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-19
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-04-19
AI Technical Summary
[0004]首先,现有技术通过对历史的光伏实际出力数据进行概率统计分析,从而得到典型出力场景,需要大量的历史数据和数据处理工作作为支撑,光伏电站采集数据类型众多,包括气候数据以及实际运行数据,历史数据处理过程复杂,计算繁琐
[0070]通过对光伏大量运行数据进行预处理,并采用基于肘形判据的K-means聚类与深度神经网络结合的双层聚类方法对光伏实测数据进行划分,最终得到典型出力场景,从而将光伏出力不确定性带来的问题转化成易于分析的问题。本发明地区光伏出力场景划分方法相较于传统新能源出力特性分析的研究方法,不再针对历史运行数据进行简单概率统计分析,本发明在典型场景聚类划分时,没有运用单一的聚类方法,而是通过改进K-means聚类算法对光伏出力历史数据先进行一次聚类,按照聚类中心对应各个k点的聚类离散度计算k点处的肘形折角,并以最小肘形折角找出目标聚类个数,即最佳聚类个数,得到一次聚类结果;再针对同类型的数据利用深度神经网络进行二次聚类,对前一次聚类结果进行反向调整修正,最终得到考虑光伏出力特性的聚类结果,从而得到结果更为精准的光伏典型出力场景,从光伏历史运行数据出发,采用了双层聚类的方法,能更科学更有效地得到光伏的典型出力场景,反映周期内光伏出力的变化特征,对含有光伏电力系统的规划和调度具有重要意义。
Smart Images

Figure CN116383688B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of planning and scheduling technology of photovoltaic power systems, specifically relating to a method, system, terminal and storage medium for dividing regional photovoltaic power output scenarios. Background Technology
[0002] With the large-scale integration of new energy sources such as wind power and solar power into the power grid, while addressing the ever-increasing load demand, it has also brought certain impacts to the operation and control of the power grid. Because solar power output is affected by factors such as location, weather, and environment, its inherent randomness and volatility pose significant challenges to the operation and control of the power grid. Since the distribution of solar resources is influenced by geographical location, seasonal climate change, and weather variations, the changes in solar power output also exhibit a certain seasonal periodicity. Characterizing the changes in solar power output within a period using typical scenarios is of great significance for the planning and dispatching of power systems containing solar power.
[0003] Existing technologies for classifying new energy power output scenarios mainly fall into two categories: one is to analyze historical actual power output data through probabilistic statistical analysis to obtain power output scenarios; the other is to use the idea of clustering, using clustering algorithms to directly cluster a large amount of new energy operation data, extract power output data with certain similarities, and analyze the clustering results to obtain power output scenarios with certain similar characteristics.
[0004] First, existing technologies derive typical output scenarios through probabilistic statistical analysis of historical photovoltaic (PV) output data. This requires substantial historical data and extensive data processing. PV power plants collect diverse data types, including climate data and actual operational data, making historical data processing complex and computationally intensive. Second, the traditional K-means algorithm randomly selects initial cluster centers, which may lead to local optima and fail to meet the global optimum, resulting in poor clustering performance. Third, artificially setting the number of clustering scenarios has drawbacks. Too many scenarios may result in minimal differences between them, while too few may lead to significant variations within each cluster, reducing the representativeness of the selected scenarios. Summary of the Invention
[0005] The purpose of this invention is to address the problems in the prior art by providing a method, system, terminal, and storage medium for classifying regional photovoltaic power output scenarios. In the case of large-scale distributed photovoltaic grid connection, the invention preprocesses a large amount of photovoltaic operation data and uses a two-layer clustering method combining improved K-means clustering and deep neural networks to classify the photovoltaic measured data, ultimately obtaining typical power output scenarios. This transforms the uncertainty brought by photovoltaics into a deterministic problem that is easy to analyze.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] Firstly, a method for classifying regional photovoltaic power output scenarios is provided, including:
[0008] K-means clustering analysis was performed on the preprocessed photovoltaic measured data. The elbow angle at each k point was calculated according to the cluster dispersion of each k point corresponding to the cluster center. The number of target clusters was found by the minimum elbow angle, and the first clustering result was obtained.
[0009] For the first clustering result, a self-organizing competitive neural network is used to obtain the second clustering result. Based on the principle of reverse correction, the second clustering result is used as the training dataset of the deep neural network to reverse correct the first clustering result. The second clustering is then repeated to obtain the final comprehensive clustering result that takes into account the photovoltaic power output characteristics, thus completing the division of the photovoltaic power output scenarios in the region.
[0010] As a preferred embodiment, in the step of performing K-means clustering analysis on the preprocessed photovoltaic measured data, the preprocessing step of the photovoltaic measured data includes filling in missing values in the data using the nearest neighbor imputation method, specifically including the following steps:
[0011] Import all known and missing data from the photovoltaic measured data;
[0012] Calculate the distance d from each missing data point to other known data points;
[0013] Sort all distances d and select the K points with the smallest distances;
[0014] Compare the categories of the selected K points and assign the sample points to be estimated to the category with the highest proportion.
[0015] As a preferred embodiment, in the step of performing K-means clustering analysis on the preprocessed photovoltaic measured data, the preprocessing step of the photovoltaic measured data includes extracting outlier data using the Isolation Forest algorithm, specifically including the following steps:
[0016] Extract n sample data from the original dataset, construct a data subset, and construct an initial binary tree iTree;
[0017] Randomly select a data feature q from all the data features to be selected, and randomly select a cut point p in the current data that is between the maximum and minimum values of the data feature q at the current node;
[0018] At the cut point p, the current data space is divided into two subspaces. Data with attribute values less than the cut point p is placed in one subspace, and data with attribute values greater than the cut point p is placed in the other subspace.
[0019] Repeatedly select a cutting point p in the current data and divide the current data space until the subspace can no longer be cut or the initial binary tree iTree has reached the initially defined limit height;
[0020] Substitute the test data x into multiple initial binary trees iTree, where h(x) represents the height of each initial binary tree iTree, E(h(x)) is the average of all h(x), H(x) is the harmonic number, and the average path length l(n) of the binary trees constructed from n samples is:
[0021] l(n) = 2H(n-1) - [2(n-1) / n]
[0022] H(i) = ln(i) + 0.5772
[0023] The anomaly score of the data to be tested is calculated using the following formula:
[0024]
[0025] The closer u(x,n) is to 1, the greater the probability that the corresponding data is abnormal data.
[0026] As a preferred embodiment, the steps of performing K-means clustering analysis on the preprocessed photovoltaic measured data, calculating the elbow angle at each k-point according to the cluster dispersion corresponding to the cluster center, and finding the target cluster size using the minimum elbow angle to obtain the first clustering result include:
[0027] k data points were randomly selected from the photovoltaic power output data as the initial cluster centers;
[0028] Data objects are assigned to the class represented by the most similar cluster center according to the Euclidean distance criterion, and the mean of all objects in each class is calculated as the new cluster center for the corresponding class.
[0029] Calculate the sum of squared distances d from all samples to the cluster centers of their respective categories. m And the value of the dispersion D(k), for any cluster m, where the distance of each sample around the cluster center is calculated by the following formula:
[0030]
[0031] In the formula, G m It is the center of cluster m;
[0032] For the number of clusters k, the dispersion is calculated using the following formula:
[0033]
[0034] In the formula, |G m |For cluster center G m The number of data points included;
[0035] Define ΔD(k) = E[ln D] ref If the increase of ΔD(k) at a certain k-value point reaches a set value, then the corresponding k-value is the target cluster value, thus obtaining a photovoltaic power output dataset with k classes having the same characteristics.
[0036] As a preferred embodiment, in the step of obtaining the secondary clustering result using a self-organizing competitive neural network based on the primary clustering result, principal component analysis is used to reduce the dimensionality of the primary clustering result data, and the dimensionality-reduced principal components are used as the input layer for the secondary clustering.
[0037] Secondly, a regional photovoltaic power output scenario division system is provided, including:
[0038] The first clustering module is used to perform K-means clustering analysis on the preprocessed photovoltaic measured data. It calculates the elbow angle at each k point according to the cluster dispersion of each k point corresponding to the cluster center, and finds the target number of clusters with the minimum elbow angle to obtain the first clustering result.
[0039] The secondary clustering module is used to obtain secondary clustering results from the primary clustering results using a self-organizing competitive neural network. Based on the principle of reverse correction, the secondary clustering results are used as the training dataset for the deep neural network to reverse correct the primary clustering results. The secondary clustering is then repeated to obtain the final comprehensive clustering result that takes into account the photovoltaic power output characteristics, thus completing the division of the photovoltaic power output scenarios in the region.
[0040] As a preferred embodiment, the regional photovoltaic output scenario division system further includes a missing value filling module, used to fill missing values in the measured photovoltaic data using the nearest neighbor interpolation method during preprocessing, specifically including:
[0041] Import all known and missing data from the photovoltaic measured data;
[0042] Calculate the distance d from each missing data point to other known data points;
[0043] Sort all distances d and select the K points with the smallest distances;
[0044] Compare the categories of the selected K points and assign the sample points to be estimated to the category with the highest proportion.
[0045] As a preferred embodiment, the regional photovoltaic output scenario segmentation system further includes an anomaly data extraction module, used to extract anomaly data from the measured photovoltaic data during preprocessing using the isolated forest algorithm, specifically including:
[0046] Extract n sample data from the original dataset, construct a data subset, and construct an initial binary tree iTree;
[0047] Randomly select a data feature q from all the data features to be selected, and randomly select a cut point p in the current data that is between the maximum and minimum values of the data feature q at the current node;
[0048] At the cut point p, the current data space is divided into two subspaces. Data with attribute values less than the cut point p is placed in one subspace, and data with attribute values greater than the cut point p is placed in the other subspace.
[0049] Repeatedly select a cutting point p in the current data and divide the current data space until the subspace can no longer be cut or the initial binary tree iTree has reached the initially defined limit height;
[0050] Substitute the test data x into multiple initial binary trees iTree, where h(x) represents the height of each initial binary tree iTree, E(h(x)) is the average of all h(x), H(x) is the harmonic number, and the average path length l(n) of the binary trees constructed from n samples is:
[0051] l(n) = 2H(n-1) - [2(n-1) / n]
[0052] H(i) = ln(i) + 0.5772
[0053] The anomaly score of the data to be tested is calculated using the following formula:
[0054]
[0055] The closer u(x,n) is to 1, the greater the probability that the corresponding data is abnormal data.
[0056] As a preferred embodiment, the primary clustering module performs K-means clustering analysis on the preprocessed photovoltaic measured data, calculates the elbow angle at each k-point according to the cluster dispersion corresponding to the cluster center, and finds the target cluster size using the minimum elbow angle to obtain the primary clustering result. The steps include:
[0057] k data points were randomly selected from the photovoltaic power output data as the initial cluster centers;
[0058] Data objects are assigned to the class represented by the most similar cluster center according to the Euclidean distance criterion, and the mean of all objects in each class is calculated as the new cluster center for the corresponding class.
[0059] Calculate the sum of squared distances d from all samples to the cluster centers of their respective categories. m And the value of the dispersion D(k), for any cluster m, where the distance of each sample around the cluster center is calculated by the following formula:
[0060]
[0061] In the formula, G m It is the center of cluster m;
[0062] For the number of clusters k, the dispersion is calculated using the following formula:
[0063]
[0064] In the formula, |G m |For cluster center G m The number of data points included;
[0065] Define ΔD(k) = E[ln D] ref If the increase of ΔD(k) at a certain k-value point reaches a set value, then the corresponding k-value is the target cluster value, thus obtaining a photovoltaic power output dataset with k classes having the same characteristics.
[0066] As a preferred embodiment, in the step of obtaining the secondary clustering result using a self-organizing competitive neural network based on the primary clustering result, the secondary clustering module employs principal component analysis to reduce the dimensionality of the primary clustering result data, and uses the dimensionality-reduced principal components as the input layer for the secondary clustering.
[0067] Thirdly, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the regional photovoltaic power output scenario division method.
[0068] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the regional photovoltaic power output scenario division method.
[0069] Compared with the prior art, the first aspect of the present invention has at least the following beneficial effects:
[0070] By preprocessing a large amount of photovoltaic operation data and using a two-layer clustering method combining K-means clustering based on the elbow criterion and deep neural networks to divide the photovoltaic measured data, typical power output scenarios are finally obtained, thereby transforming the problem caused by the uncertainty of photovoltaic power output into a problem that is easy to analyze. Compared to traditional research methods for analyzing the characteristics of new energy power output, the method for classifying photovoltaic (PV) power output scenarios in this invention does not rely on simple probabilistic statistical analysis of historical operating data. Instead of using a single clustering method, this invention employs an improved K-means clustering algorithm to first cluster historical PV power output data. The elbow angle at each k-point is calculated based on the cluster dispersion corresponding to the cluster center, and the optimal number of clusters (the minimum elbow angle) is used to determine the first clustering result. Then, a second clustering is performed using a deep neural network on similar data, adjusting and correcting the previous clustering result. Finally, a clustering result considering PV power output characteristics is obtained, leading to a more accurate representation of typical PV power output scenarios. Starting from historical PV operating data, this two-layer clustering method more scientifically and effectively identifies typical PV power output scenarios, reflecting the changing characteristics of PV power output within a cycle. This is of great significance for the planning and scheduling of PV-containing power systems.
[0071] It is understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0072] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0073] Figure 1 Flowchart of the method for dividing regional photovoltaic power output scenarios according to an embodiment of the present invention;
[0074] Figure 2 The embodiment of this invention is based on a two-layer clustering principle diagram combining K-means clustering and deep neural networks;
[0075] Figure 3 A structural block diagram of a regional photovoltaic power output scenario division system according to an embodiment of the present invention. Detailed Implementation
[0076] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0077] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0078] Example 1
[0079] The method for classifying regional photovoltaic (PV) output scenarios in this invention addresses the situation of large-scale distributed PV grid connection. It preprocesses a large amount of PV operational data and employs a two-layer clustering method combining improved K-means clustering and deep neural networks to classify the measured PV data, ultimately obtaining typical output scenarios. This transforms the uncertainties brought by PV into deterministic problems that are easy to analyze. Please refer to... Figure 1 and Figure 2 The embodiments of the present invention specifically include the following steps:
[0080] S1. Handling Missing Data in Photovoltaic Measured Data
[0081] During the collection of photovoltaic measured data, data loss can occur due to equipment malfunctions or human error. The K-Nearest Neighbor (KNN) method is used to fill in missing values. The core idea is simple: if a sample has K nearest neighbors belonging to the same category in a feature space, then that sample belongs to that category and possesses the characteristics of samples in that category. The KNN algorithm is well-suited for automatic classification with large sample sizes and for class classification problems where there is overlap between class domains. Therefore, it is suitable for filling in missing data in distributed renewable energy output. The missing data processing in this embodiment mainly consists of the following four steps:
[0082] Step 1.1: Import all known data and missing data;
[0083] Step 1.2: Calculate the distance d from each missing data point to other known data points;
[0084] Step 1.3: Sort all distances d and select the K points with the smallest distances;
[0085] Step 1.4: Compare the K categories selected above and assign the sample point to be estimated to the category with the highest proportion.
[0086] S2. Handling Abnormal Data in Photovoltaic Measured Data
[0087] Since most photovoltaic power plants are built in remote areas with abundant photovoltaic resources, far from the monitoring center, high-sensitivity sensors are highly susceptible to adverse weather conditions during the data acquisition phase, generating a large amount of unstable data that deviates from the original actual values, thus producing anomalies. The Isolation Forest algorithm is used to extract outliers from the data. This algorithm has linear time complexity and high accuracy, low computational cost, and is very suitable for massive data scenarios, meeting the requirements of big data processing. The Isolation Forest algorithm considers sparse data points far from denser clusters as outliers, and is suitable for extracting anomalies in photovoltaic output data caused by equipment problems or human factors. The steps include:
[0088] (1) Extract n sample data from the original dataset with or without replacement, construct a data subset, and construct an initial binary tree iTree.
[0089] (2) Randomly select a data feature q from all the data to be selected, and randomly select a cutting point p in the current data that is between the maximum and minimum values of the current node's data feature q.
[0090] (3) Divide the current data space into two subspaces, left and right, at the cut point p. Place the attribute values of the data items that are less than the cut point p into the left subspace, and place the attribute values of the data items that are greater than the cut point p into the right subspace.
[0091] (4) Repeat steps (2) and (3) until the subspace can no longer be cut or the initial binary tree iTree has reached the initially defined limit height.
[0092] For the test data x, substitute it into each initial binary tree iTree, let h(x) represent its height on each initial binary tree iTree, E(h(x)) be the average of all h(x), and H(x) be the harmonic number. Because the structure of the initial binary tree iTree is similar to that of a binary search tree, the average path length l(n) for constructing a binary tree for n samples is:
[0093] l(n) = 2H(n-1) - [2(n-1) / n]
[0094] H(i) = ln(i) + 0.5772
[0095] Define an anomaly score for a test data point as:
[0096]
[0097] The closer u(x,n) is to 1, the greater the probability that the data is outlier.
[0098] S3. Improved K-means clustering based on GSA elbow criterion
[0099] After processing the measured photovoltaic data, clustering is necessary to further explore the characteristics of regional photovoltaic output. Clustering is a crucial analytical method in data mining. This invention proposes a two-layer clustering method combining an improved K-means clustering method based on the GSA elbow criterion and a deep neural network. The specific implementation process is as follows:
[0100] 1) First, randomly select k data points from the photovoltaic power output data as initial cluster centers;
[0101] 2) Assign data objects to the class represented by the cluster center that is most similar to them according to the Euclidean distance criterion, and calculate the mean of all objects in each class as the new cluster center of that class;
[0102] 3) Calculate the sum of squared distances d from all samples to the cluster centers of their respective categories. m And the value of the dispersion D(k). For any cluster m, where the distance of each sample around the cluster center is:
[0103]
[0104] In the formula, G m It is the center of cluster m.
[0105] For the number of clusters k, its dispersion can be expressed as:
[0106]
[0107] In the formula, |G m |For cluster center G m The number of data points included.
[0108] Define αD(k) = E[ln D] ref If ΔD(k) increases significantly at a certain k value, then the corresponding k value is the optimal number of clusters, resulting in a photovoltaic power output dataset with k classes having the same characteristics.
[0109] S4. Quadratic clustering based on deep neural networks
[0110] For similar photovoltaic measured data, a secondary clustering method is performed. Principal component analysis is used to reduce the dimensionality of the data features, and the dimensionality-reduced principal components are used as the input layer for the secondary clustering. The pattern classification of a self-organizing competitive neural network is then used to obtain the secondary clustering results. Then, based on the principle of back-correction, the secondary clustering results are used as the training dataset for a deep neural network to back-correct the primary clustering results. The secondary clustering process is repeated to obtain the final comprehensive clustering result that considers the photovoltaic power output characteristics.
[0111] First, compared with traditional new energy output scenario classification methods, the regional photovoltaic output scenario classification method of this invention no longer performs simple probability statistical analysis on historical operating data. Second, when clustering typical scenarios, the regional photovoltaic output scenario classification method of this invention does not use a single clustering method. Instead, it uses an improved K-means clustering algorithm to first cluster the historical photovoltaic output data, and then uses a deep neural network to perform a second clustering on the same type of data. The results of the first clustering are adjusted and corrected in reverse, and finally a clustering result that takes into account the characteristics of photovoltaic output is obtained, thus obtaining a more accurate result of typical photovoltaic output scenarios.
[0112] Example 2
[0113] Please see Figure 3 This invention also proposes a regional photovoltaic power output scenario division system, including:
[0114] The first clustering module is used to perform K-means clustering analysis on the preprocessed photovoltaic measured data. It calculates the elbow angle at each k point according to the cluster dispersion of each k point corresponding to the cluster center, and finds the target number of clusters with the minimum elbow angle to obtain the first clustering result.
[0115] The secondary clustering module is used to obtain secondary clustering results from the primary clustering results using a self-organizing competitive neural network. Based on the principle of reverse correction, the secondary clustering results are used as the training dataset for the deep neural network to reverse correct the primary clustering results. The secondary clustering is then repeated to obtain the final comprehensive clustering result that takes into account the photovoltaic power output characteristics, thus completing the division of the photovoltaic power output scenarios in the region.
[0116] In one possible implementation, the regional photovoltaic output scenario division system of this invention further includes a missing value filling module, used to fill missing values in the photovoltaic measured data using the nearest neighbor interpolation method during preprocessing, specifically including:
[0117] Import all known and missing data from the photovoltaic measured data;
[0118] Calculate the distance d from each missing data point to other known data points;
[0119] Sort all distances d and select the K points with the smallest distances;
[0120] Compare the categories of the selected K points and assign the sample points to be estimated to the category with the highest proportion.
[0121] In one possible implementation, the regional photovoltaic output scenario segmentation system of this invention further includes an anomaly data extraction module, used to extract anomaly data from the measured photovoltaic data during preprocessing using the isolated forest algorithm, specifically including:
[0122] Extract n sample data from the original dataset, construct a data subset, and construct an initial binary tree iTree;
[0123] Randomly select a data feature q from all the data features to be selected, and randomly select a cut point p in the current data that is between the maximum and minimum values of the data feature q at the current node;
[0124] At the cut point p, the current data space is divided into two subspaces. Data with attribute values less than the cut point p is placed in one subspace, and data with attribute values greater than the cut point p is placed in the other subspace.
[0125] Repeatedly select a cutting point p in the current data and divide the current data space until the subspace can no longer be cut or the initial binary tree iTree has reached the initially defined limit height;
[0126] Substitute the test data x into multiple initial binary trees iTree, where h(x) represents the height of each initial binary tree iTree, E(h(x)) is the average of all h(x), H(x) is the harmonic number, and the average path length l(n) of the binary trees constructed from n samples is:
[0127] l(n) = 2H(n-1) - [2(n-1) / n]
[0128] H(i) = ln(i) + 0.5772
[0129] The anomaly score of the data to be tested is calculated using the following formula:
[0130]
[0131] The closer u(x,n) is to 1, the greater the probability that the corresponding data is abnormal data.
[0132] In one possible implementation, the first-order clustering module of this embodiment performs K-means clustering analysis on the preprocessed photovoltaic measured data, calculates the elbow angle at each k-point according to the cluster dispersion corresponding to the cluster center, and finds the target cluster number using the minimum elbow angle to obtain the first-order clustering result.
[0133] k data points were randomly selected from the photovoltaic power output data as the initial cluster centers;
[0134] Data objects are assigned to the class represented by the most similar cluster center according to the Euclidean distance criterion, and the mean of all objects in each class is calculated as the new cluster center for the corresponding class.
[0135] Calculate the sum of squared distances d from all samples to the cluster centers of their respective categories. m And the value of the dispersion D(k), for any cluster m, where the distance of each sample around the cluster center is calculated by the following formula:
[0136]
[0137] In the formula, G m It is the center of cluster m;
[0138] For the number of clusters k, the dispersion is calculated using the following formula:
[0139]
[0140] In the formula, |G m |For cluster center G m The number of data points included;
[0141] Define ΔD(k) = E[ln D] ref If the increase of ΔD(k) at a certain k-value point reaches a set value, then the corresponding k-value is the target cluster value, thus obtaining a photovoltaic power output dataset with k classes having the same characteristics.
[0142] In one possible implementation, in the step of obtaining secondary clustering results using a self-organizing competitive neural network based on the primary clustering results, the secondary clustering module of this embodiment of the invention uses principal component analysis to reduce the dimensionality of the data from the primary clustering results, and uses the dimensionality-reduced principal components as the input layer for secondary clustering.
[0143] Example 3
[0144] Embodiments of the present invention also propose an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the regional photovoltaic power output scenario division method.
[0145] Example 4
[0146] Embodiments of the present invention also propose a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the regional photovoltaic power output scenario division method.
[0147] The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable storage medium can include any entity or device capable of carrying the computer program code, a medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory, a random access memory, an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals. For ease of explanation, the above content only shows the parts related to the embodiments of the present invention; for specific technical details not disclosed, please refer to the method section of the embodiments of the present invention. This computer-readable storage medium is non-transitory and can be stored in storage devices formed by various electronic devices, enabling the execution process described in the method of the embodiments of the present invention.
[0148] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0149] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0150] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0151] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0152] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for dividing regional photovoltaic power output scenarios, characterized in that, include: K-means clustering analysis was performed on the preprocessed photovoltaic measured data. The elbow angle at each k point was calculated according to the cluster dispersion of each k point corresponding to the cluster center. The number of target clusters was found by the minimum elbow angle, and the first clustering result was obtained. For the results of the first clustering, a self-organizing competitive neural network is used to obtain the results of the second clustering. Based on the principle of reverse correction, the secondary clustering results are used as the training dataset for the deep neural network to reverse correct the primary clustering results. The secondary clustering is then repeated to obtain the final comprehensive clustering results that take into account the photovoltaic power output characteristics, thus completing the division of the photovoltaic power output scenarios in the region. In the step of obtaining the secondary clustering results using a self-organizing competitive neural network based on the primary clustering results, principal component analysis is used to reduce the dimensionality of the primary clustering data, and the dimensionality-reduced principal components are used as the input layer for the secondary clustering. The steps of performing K-means clustering analysis on the preprocessed photovoltaic measured data, calculating the elbow angle at each k-point according to the cluster dispersion corresponding to the cluster center, and finding the target cluster size using the minimum elbow angle to obtain the first clustering result include: k data points were randomly selected from the photovoltaic power output data as the initial cluster centers; Data objects are assigned to the class represented by the most similar cluster center according to the Euclidean distance criterion, and the mean of all objects in each class is calculated as the new cluster center for the corresponding class. Calculate the sum of squared distances from all samples to the cluster centers of their respective categories. and dispersion The value of is calculated for any cluster m, where the distance of each sample around the cluster center is calculated as follows: In the formula, It is the center of cluster m; For the number of clusters k, the dispersion is calculated using the following formula: In the formula, Cluster center The number of data points included; definition If at a certain k value point When the increase of k reaches the set value, the corresponding k value is the target cluster number, thus obtaining a photovoltaic power output dataset with k classes having the same characteristics.
2. The method for dividing regional photovoltaic power output scenarios according to claim 1, characterized in that, In the step of performing K-means clustering analysis on the preprocessed photovoltaic measured data, the preprocessing step includes filling in missing values in the data using the nearest neighbor imputation method, specifically including the following steps: Import all known and missing data from the photovoltaic measured data; Calculate the distance d from each missing data point to other known data points; Sort all distances d and select the K points with the smallest distances; Compare the categories of the selected K points and assign the sample points to be estimated to the category with the highest proportion.
3. The method for dividing regional photovoltaic power output scenarios according to claim 1, characterized in that, In the step of performing K-means clustering analysis on the preprocessed photovoltaic measured data, the preprocessing step includes using the isolated forest algorithm to extract outlier data from the data, specifically including the following steps: Extract n sample data from the original dataset, construct a data subset, and construct an initial binary tree iTree; Randomly select a data feature q from all the data features to be selected, and randomly select a cut point p in the current data that is between the maximum and minimum values of the data feature q at the current node; At the cut point p, the current data space is divided into two subspaces. Data with attribute values less than the cut point p is placed in one subspace, and data with attribute values greater than the cut point p is placed in the other subspace. Repeatedly select a cutting point p in the current data and divide the current data space until the subspace can no longer be cut or the initial binary tree iTree has reached the initially defined limit height; Substitute the data to be tested x into multiple initial binary trees iTree. The height of each initial binary tree iTree represents the height of the tree. For all The average value, Let n be the harmonic number, the average path length of the binary tree constructed from n samples. for: The anomaly score of the data to be tested is calculated using the following formula: In the formula The closer the value is to 1, the greater the probability that the corresponding data is abnormal.
4. A regional photovoltaic power output scenario division system, characterized in that, include: The first clustering module is used to perform K-means clustering analysis on the preprocessed photovoltaic measured data. It calculates the elbow angle at each k point according to the cluster dispersion of each k point corresponding to the cluster center, and finds the target number of clusters with the minimum elbow angle to obtain the first clustering result. The secondary clustering module is used to obtain secondary clustering results from the primary clustering results using a self-organizing competitive neural network; Based on the principle of reverse correction, the secondary clustering results are used as the training dataset for the deep neural network to reverse correct the primary clustering results. The secondary clustering is then repeated to obtain the final comprehensive clustering result that takes into account the photovoltaic power output characteristics, thus completing the division of the photovoltaic power output scenarios in the region. In the step of obtaining the secondary clustering results using a self-organizing competitive neural network based on the primary clustering results, the secondary clustering module uses principal component analysis to reduce the dimensionality of the primary clustering data and uses the dimensionality-reduced principal components as the input layer for the secondary clustering. The first-order clustering module performs K-means clustering analysis on the preprocessed photovoltaic measured data, calculates the elbow angle at each k-point according to the cluster dispersion corresponding to the cluster center, and finds the target cluster size by using the minimum elbow angle to obtain the first-order clustering result. The steps include: k data points were randomly selected from the photovoltaic power output data as the initial cluster centers; Data objects are assigned to the class represented by the most similar cluster center according to the Euclidean distance criterion, and the mean of all objects in each class is calculated as the new cluster center for the corresponding class. Calculate the sum of squared distances from all samples to the cluster centers of their respective categories. and dispersion The value of is calculated for any cluster m, where the distance of each sample around the cluster center is calculated as follows: In the formula, It is the center of cluster m; For the number of clusters k, the dispersion is calculated using the following formula: In the formula, Cluster center The number of data points included; definition If at a certain k value point When the increase of k reaches the set value, the corresponding k value is the target cluster number, thus obtaining a photovoltaic power output dataset with k classes having the same characteristics.
5. The regional photovoltaic power output scenario division system according to claim 4, characterized in that, It also includes a missing value filling module, which is used to fill missing values in photovoltaic measured data using the nearest neighbor interpolation method during preprocessing. Specifically, it includes: Import all known and missing data from the photovoltaic measured data; Calculate the distance d from each missing data point to other known data points; Sort all distances d and select the K points with the smallest distances; Compare the categories of the selected K points and assign the sample points to be estimated to the category with the highest proportion.
6. The regional photovoltaic power output scenario division system according to claim 4, characterized in that, It also includes an anomaly data extraction module, used during preprocessing to extract anomalies from measured photovoltaic data using the Isolation Forest algorithm, specifically including: Extract n sample data from the original dataset, construct a data subset, and construct an initial binary tree iTree; Randomly select a data feature q from all the data features to be selected, and randomly select a cut point p in the current data that is between the maximum and minimum values of the data feature q at the current node; At the cut point p, the current data space is divided into two subspaces. Data with attribute values less than the cut point p is placed in one subspace, and data with attribute values greater than the cut point p is placed in the other subspace. Repeatedly select a cutting point p in the current data and divide the current data space until the subspace can no longer be cut or the initial binary tree iTree has reached the initially defined limit height; Substitute the data to be tested x into multiple initial binary trees iTree. The height of each initial binary tree iTree represents the height of the tree. For all The average value, Let n be the harmonic number, the average path length of the binary tree constructed from n samples. for: The anomaly score of the data to be tested is calculated using the following formula: In the formula The closer the value is to 1, the greater the probability that the corresponding data is abnormal.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the steps of the regional photovoltaic power output scenario division method as described in any one of claims 1 to 3.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by the processor, it implements the steps of the regional photovoltaic power output scenario division method as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Distributed photovoltaic multi-scene analysis method based on H-K composite clustering algorithm
CN111695586A
Double-layer adaptive clustering method considering load characteristics and adjustable potential
CN115861671A