Small hydropower station power generation data abnormal value identification and correction method and system
By identifying and correcting outliers and vacant values in small hydropower generation data, data quality problems are solved, data integrity and accuracy are improved, and reliable data support is provided for power grid scheduling.
Patent Information
- Application Number
- CN202510014106.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-06-10
AI Technical Summary
There are outliers and vacant values in the small hydropower generation data, which leads to errors in the prediction of power generation capacity and affects the accuracy and efficiency of power grid scheduling.
By replacing the zero value in the small hydropower generation dataset with a null value, the data rate of change characteristics are extracted, outliers are identified based on the normal distribution theory and empty them, and a waveform library is constructed to fill the vacant values.
It significantly improves the integrity and accuracy of data, avoids human error, ensures data reliability, and provides solid data support for grid scheduling and load prediction.
Smart Images

Figure CN120123916A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system analysis, and particularly to a method and system for identifying and correcting outliers in small hydropower generation data. Background Art
[0002] Due to its small scale, low investment, and quick results, small hydropower has become an important force in promoting local economic development. Nevertheless, small hydropower resources are usually distributed in remote and relatively economically backward areas, where the electricity demand is relatively limited, resulting in the generation capacity often far exceeding the demand, causing a waste of power resources. Especially during the flood season, due to the synchrony of water sources, large and small hydropower systems may simultaneously be at the peak of power generation, causing tension in the transmission channel resources, further leading to power curtailment and water abandonment. This waste of resources not only reduces the power utilization efficiency but also may pose a threat to the stability of the power grid. In the power generation prediction of small hydropower, data quality is crucial. Since the power generation process is affected by various factors, such as water flow, meteorological conditions, historical power generation records, etc., the collected data may contain a certain degree of noise or outliers. These outliers may stem from various reasons, such as instrument failures, data entry errors, extreme weather, etc. If these outliers are not identified and corrected in a timely manner, it will lead to errors in the prediction of power generation capacity, thus affecting the accuracy and efficiency of power grid dispatching.
[0003] If outliers are not processed in a timely manner, it will lead to the distortion of data analysis results. For example, if the historical power generation data of a small hydropower station shows abnormal fluctuations due to a fault, and this fluctuation is misinterpreted as a normal change, it may cause the system to wrongly dispatch the power grid, resulting in over-reliance on certain power sources or ignoring the actual existing power generation capacity. This may not only exacerbate the problems of power curtailment and water abandonment but also lead to power shortages or surpluses, affecting the stable operation of the power grid.
[0004] The power generation prediction of small hydropower depends on accurate input data, including water flow, rainfall, temperature, historical power generation conditions, etc. If there are outliers in the input data, the accuracy of the prediction model will be severely affected. The existence of outliers may cause the model to deviate from the actual situation, generating overly high or low prediction values, thus affecting the power grid dispatching decision-making and reducing the utilization efficiency of power resources. Therefore, how to optimize the dispatching of the power system and improve the scientific nature of resource allocation is the key to solving this problem. Summary of the Invention
[0005] In view of the problems existing in the above-mentioned prior art, the present invention is proposed.
[0006] Therefore, the problems to be solved by the present invention are how to identify outliers in the small hydropower generation data of a region and how to fill in the missing values.
[0007] To solve the above technical problems, in a first aspect, the present invention provides the following technical solution: A method for identifying and correcting outliers in small hydropower generation data, which includes replacing zero values in the small hydropower generation data set with null values; extracting data change rate features based on the change relationship between adjacent data points in the small hydropower generation data set; identifying outliers in the small hydropower generation data set based on the normal distribution theory and performing nulling processing; extracting continuous normal data segments in the small hydropower generation data set and constructing a waveform library; filling in the missing values in the small hydropower generation data set according to the normal data samples in the waveform library.
[0008] As a preferred solution of the method for identifying and correcting outliers in small hydropower generation data according to the present invention, wherein: the data change rate features include the forward power change rate D forw (t) and the backward power change rate D back (t), and their expressions are respectively:
[0009]
[0010] wherein, load(t) represents the load data at time t.
[0011] As a preferred solution of the method for identifying and correcting outliers in small hydropower generation data according to the present invention, wherein: the relevant parameters of the normal distribution include the mean μ forw of the forward power change rate D forw , the standard deviation σ forw of the forward power change rate D forw (t), the mean μ back of the backward power change rate D back and the standard deviation σ back of the backward power change rate D back (t), and the expressions are respectively:
[0012]
[0013] As a preferred solution of the method for identifying and correcting outliers in small hydropower generation data according to the present invention, wherein: the method for identifying outliers in the small hydropower generation data set includes comparing the forward power change rate D forw (t) and the backward power change rate D back (t) with the confidence intervals defined by their respective means and standard deviations respectively. If the change rate of the data point exceeds the confidence interval, it is regarded as an outlier and nulled.
[0014] As a preferred solution of the method for identifying and correcting outliers in small hydropower generation data according to the present invention, wherein: the forward power change rate D forwThe confidence interval of (t) is [μ forw -4σ forw ,μ forw +4σ forw . The confidence interval of the backward power change rate D back (t) is [μ back -lσ back , μ back +lσ back , where l is an adjustment parameter used to adjust the width of the confidence interval.
[0015] As a preferred solution of the method for identifying and correcting outliers in small hydropower generation data according to the present invention, wherein: the method for constructing the waveform library includes separating from the null values in the small hydropower generation dataset, screening out all continuous normal data segments, and sampling using the sliding window technique.
[0016] As a preferred solution of the method for identifying and correcting outliers in small hydropower generation data according to the present invention, wherein: the size of the sliding window is 48 and the sliding step is 1.
[0017] As a preferred solution of the method for identifying and correcting outliers in small hydropower generation data according to the present invention, wherein: the KNN algorithm is used to fill the missing values in the small hydropower generation dataset. The method includes calculating the distance between the data point to be filled and each waveform segment in the waveform library, finding k nearest waveform segments, and filling the missing values according to the weighted average of these neighboring waveform segments.
[0018] As a preferred solution of the method for identifying and correcting outliers in small hydropower generation data according to the present invention, wherein: the KNN algorithm calculates the distance between the sequence to be filled and the normal waveform segments in the waveform library through the Euclidean distance, calculates the weighting coefficient according to the reciprocal of the distance, and the waveform segments with closer distances obtain greater weights. Finally, the missing data is filled by the weighted average.
[0019] On the other hand, the present invention also provides a system for identifying and correcting outliers in small hydropower generation data, which is applicable to the above method for identifying and correcting outliers in small hydropower generation data. The system includes a data preprocessing module for traversing the small hydropower generation dataset and replacing the zero values therein with null values; a data change rate calculation module for calculating the forward power change rate and the backward power change rate; an outlier identification module for setting a confidence interval according to the mean and standard deviation of the forward and backward change rates, identifying the outliers beyond the confidence interval and performing nulling processing; and a data filling module for obtaining normal data samples from the waveform library and filling the missing values in the dataset.
[0020] The beneficial effects of the present invention are as follows: By effectively identifying and correcting zero values, outliers, and missing values in the regional small hydropower generation data, the integrity and accuracy of the data are significantly improved. The outlier identification algorithm based on the data change rate and the KNN filling method are adopted, which can not only automatically process missing data but also avoid human errors, ensuring the reliability of the data, providing solid data support for subsequent power grid dispatching, load forecasting, and optimization decision-making, and enhancing the intelligent level and operation efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0022] Figure 1 It is a schematic diagram of the overall steps of the outlier identification and correction method for small hydropower generation data.
[0023] Figure 2 It is a schematic diagram of the detailed steps of the outlier identification and correction method for small hydropower generation data. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings of the specification.
[0025] In the following description, many specific details are set forth to facilitate a thorough understanding of the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0026] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it an embodiment that is separately or selectively mutually exclusive with other embodiments.
[0027] Embodiment 1
[0028] Referring to Figure 1 and Figure 2 , this is the first embodiment of the present invention. This embodiment provides an outlier identification and correction method for small hydropower generation data, and the method includes:
[0029] S1: Replace the zero values in the small hydropower generation data set with null values.
[0030] In many datasets, zero values are often used to represent missing or invalid data. For example, if a sensor fails to read data during data collection, the system may record it as zero.
[0031] If there are a large number of zero values in the data, and these zero values do not actually represent real power or other meaningful values, then they will lead to inaccurate subsequent analysis. Therefore, replacing zero values with null values first can help us identify these invalid data.
[0032] In small hydropower generation data, the power generation should not be zero (except in some special cases, such as complete shutdown or failure). Generally, the power data should be a value greater than zero.
[0033] If zero values appear in the data, it is usually because of errors in data input or collection. These zero values do not represent the actual power generation status, so they need to be cleared or replaced with null values to avoid affecting subsequent analysis and processing.
[0034] Therefore, traverse the small hydropower generation dataset through a program and check all data points one by one. For each data point, check whether its value is zero. If the value of a data point is zero, it is considered an invalid data (such as data loss or incorrect entry). When a data point with a zero value is found, replace the value of that data point with a null value (which can be an empty string, NaN, or null, depending on the programming language or data storage specification).
[0035] The program continues to check other data points in the dataset until all zero values in the entire dataset are replaced with null values.
[0036] S2: Extract data change rate features based on the change relationship between adjacent data points in the small hydropower generation dataset.
[0037] Specifically, in small hydropower data, although the data curve will show strong non - linearity, the data change rate between any two adjacent data points will be relatively stable. Therefore, calculate the forward power change rate D forw (t) and the backward power change rate D back (t) of the data at time t compared to the adjacent time, and their expressions are respectively:
[0038]
[0039] Among them, load(t) represents the load data at time t.
[0040] The forward power change rate D forw (t) and the backward power change rate D back(t) is used to capture the changing trend in the data. By calculating the load difference between the current moment and the next moment, it is possible to determine whether the data is changing normally or if there are abnormal fluctuations. This method of calculating the rate of change helps analyze the stability of the load change during the power generation process in order to identify possible abnormal points or emergencies.
[0041] S3: Based on the normal distribution theory, identify the outliers in the small hydropower generation dataset and perform nulling processing.
[0042] The normal distribution (also known as the Gaussian distribution) is a symmetric, bell-shaped probability distribution. In this embodiment, the normal distribution theory is used to model the rate of change in the small hydropower generation dataset and help identify outliers. The specific application steps are as follows:
[0043] S3-1: Calculate the statistical parameters of the rate of change.
[0044] First, in S2, the forward power change rate D forw (t) and the backward power change rate D back (t) were calculated. These rates of change usually exhibit a relatively stable distribution, so it can be assumed that they conform to the normal distribution. Among them:
[0045] The mean μ forw of the forward power change rate D forw : is the average value of all data points of the forward power change rate.
[0046] The standard deviation σ forw of the forward power change rate D forw : is the degree of dispersion of the forward power change rate data points.
[0047] The mean μ back of the backward power change rate D back : is the average value of all data points of the backward power change rate.
[0048] The standard deviation σ back of the backward power change rate D back : is the degree of dispersion of the backward power change rate data points.
[0049] These statistical parameters determine the normal fluctuation range of the data by calculating the mean and standard deviation of the normal distribution.
[0050] S3-2: Identify outliers.
[0051] According to the normal distribution theory, the rate of change of normal data should be near the mean, and the fluctuation of the rate of change does not exceed a certain range. Therefore, based on the mean and standard deviation of the forward and backward power change rates, a confidence interval can be set to determine which data points are abnormal, that is:
[0052] Forward power change rate D forw (t)'s confidence interval: Set to [μ forw -4σ forw , μ forw +4σ forw , that is, if the value of the forward power change rate D forw (t) exceeds this range, then this data point may be considered abnormal.
[0053] Backward power change rate D back (t): Set to [μ back -lσ back , μ back +lσ back , where l is an adjustment coefficient used to adjust the confidence interval of the backward change rate.
[0054] If the forward and backward change rates of a certain data point exceed these two confidence intervals at the same time, then this data point is considered an outlier and will be set to a null value.
[0055] S4: Extract continuous normal data segments from the small hydropower generation dataset and construct a waveform library.
[0056] For each selected continuous normal data segment, next apply the sliding window technique to extract data.
[0057] Set a window size of 48 (for example: each sliding window contains 48 data points), and set the sliding step size to 1 (that is, slide one data point each time).
[0058] For example, a certain continuous data in the small hydropower generation data is:
[0059] [x 1 , x 2 , x 3 ,......, x 100
[0060] Use the sliding window method to extract different small waveform segments from it:
[0061] First window: [x 1 , x 2 ,..., x 48
[0062] Second window: [x 2 , x 3 ,..., x 49
[0063] Third window: [x 3 , x4 ,...,x 50
[0064] And so on until the entire data segment is traversed by the window.
[0065] In this way, each window will extract a small segment of data to generate a waveform segment.
[0066] Waveform segments are extracted from multiple normal data segments through a sliding window and finally aggregated into a waveform library. This waveform library contains waveform segments extracted from all normal data segments, and each waveform segment has a length of 48 data points. Each waveform segment in the waveform library can be regarded as a feature, and these waveform segments represent the power generation power change pattern during normal operation.
[0067] S5: Fill in the missing values in the small hydropower generation dataset according to the normal data samples in the waveform library.
[0068] S5-1: Define the sequence to be filled.
[0069] Suppose the power generation data at a certain time point in the dataset is missing and needs to be filled. The input of the KNN algorithm is the data at this time point and the previous 47 time points (a total of 48 data points), and this data segment is called the sequence to be filled. Suppose the value of this sequence to be filled is:
[0070] X = [x 1 , x 2 ,..., x 47 , Z]
[0071] Among them, Z represents the missing value to be filled.
[0072] S5-2: Calculate the distance.
[0073] For the sequence to be filled X, it is necessary to compare it with each normal waveform segment in the waveform library and calculate the similarity between them. The distance calculation method uses the Euclidean distance. Suppose a certain normal waveform segment in the waveform library is Y k , and the data points it contains are:
[0074] Y k = [y 1 , y 2 ,..., y 47 , y 48
[0075] The KNN algorithm will calculate the distance between the sequence to be filled and each normal waveform segment in the waveform library. The calculation formula for the distance is:
[0076]
[0077] S5-3: Select k neighboring sequences.
[0078] By calculating the distances between the sequence X to be filled and all normal waveform segments in the waveform library, the similarity (i.e., distance) between each waveform segment and the sequence X to be filled is obtained. Then, the KNN algorithm selects k waveform segments with the closest distances (i.e., the k waveform segments most similar to the sequence X to be filled).
[0079] S5-4: Weighted average.
[0080] For these k neighboring sequences Y k , the KNN algorithm weights them according to the distances between them and the sequence X to be filled, and calculates their weighted coefficients. Sequences with closer distances are given greater weights.
[0081] The way to calculate the weighted coefficient is to use the reciprocal of the distance (the closer the neighbor, the greater the weight):
[0082]
[0083] where d(X, Y k ) is the distance between the sequence X to be filled and the waveform segment Y k .
[0084] Then, the weighted average is calculated to obtain the filling value at the position of the missing value in the sequence X to be filled. Assuming the position of the missing value is X p , the calculation formula for the filling value is as follows:
[0085]
[0086] where Y k (p) represents the value of the k-th neighboring sequence at the missing position p.
[0087] S5-5: Fill in the missing values.
[0088] According to the above weighted average calculation result, the KNN algorithm uses the calculated filling value to fill in the missing values in the sequence X to be filled. Finally, the missing values of the data points to be filled are replaced with the most appropriate values.
[0089] Embodiment 2
[0090] This embodiment is the second embodiment of the present invention, and this embodiment is based on the previous embodiment. This embodiment provides a small hydropower generation data outlier identification and correction system, which is applicable to the above-mentioned small hydropower generation data outlier identification and correction method. This system can automatically process and correct zero values, outliers, and missing values in the data, providing high-quality data support for the power generation monitoring system.
[0091] The system includes a data preprocessing module, a data change rate calculation module, an outlier identification module, and a data filling module.
[0092] Specifically, the input of the data preprocessing module is the small hydropower generation dataset. This module receives the power generation data and identifies the missing data by judging the zero values in the dataset, replacing the zero values with null values (NA). For example, if the power generation load at a certain moment is zero, indicating data loss, it is marked as NA.
[0093] The data change rate calculation module inputs the processed power generation dataset. This module calculates the forward power change rate D forw (t) and the backward power change rate D back (t). Finally, it outputs the calculation results of the forward power change rate D forw (t) and the backward power change rate D back (t).
[0094] The outlier identification module inputs the forward and backward power change rate data, the calculated mean and standard deviation. Based on the normal distribution theory, this module calculates the mean and standard deviation of the forward and backward change rates, and then sets the confidence intervals:
[0095] Forward power change rate confidence interval: [μ forw -4σ forw , μ forw +4σ forw .
[0096] Backward power change rate confidence interval: [μ back -lσ back , μ back +lσ back .
[0097] Identify the outliers beyond these confidence intervals and mark them as null values.
[0098] The data filling module inputs the power generation dataset with marked outliers, the waveform library, and the data sequence to be filled. This module uses the KNN algorithm to calculate the distance between the data to be filled and the normal data segments in the waveform library, and selects the K most similar data segments. Then, according to the weighted average of these neighboring segments, it fills the missing values in the dataset.
[0099] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for identifying and correcting abnormal values in small hydropower generation data, characterized by: include, Replace zero values in the small hydropower generation dataset with null values; Extracting data change rate features based on the change relationship between adjacent data points in the small hydropower generation data set; Based on the normal distribution theory, outliers in the small hydropower generation data set are identified and removed; Extracting continuous normal data segments from the small hydropower generation data set and constructing a waveform library; The missing values in the small hydropower generation data set are filled according to the normal data samples in the waveform library.
2. The method for identifying and correcting abnormal values of small hydropower generation data according to claim 1, characterized in that: The data change rate characteristics include the forward power change rate D forw (t) and the backward power change rate D back (t), their expressions are: Wherein, load(t) represents the load data at time t.
3. The method for identifying and correcting abnormal values of small hydropower generation data according to claim 2, characterized in that: The related parameters of the normal distribution include the forward power change rate D forw The mean value μ of (t) forw , the forward power change rate D forw The standard deviation of (t) forw , the backward power change rate D back The mean value μ of (t) back And the backward power change rate D back The standard deviation of (t) back , the expressions are:
4. The method for identifying and correcting abnormal values of small hydropower generation data according to claim 3, characterized in that: The method for identifying outliers in the small hydropower generation data set includes: forw (t) and the backward power change rate D back (t) are compared with the confidence intervals defined by their respective means and standard deviations. If the rate of change of a data point exceeds the confidence interval, it is considered an outlier and is set to blank.
5. The method for identifying and correcting abnormal values of small hydropower generation data according to claim 4, characterized in that: The forward power change rate D forw The confidence interval of (t) is [μ forw -4σ forw , μ forw +4σ forw ], the backward power change rate D back The confidence interval of (t) is [μ back -lσ back ,μ back +lσ back ], where l is an adjustment parameter used to adjust the width of the confidence interval.
6. The method for identifying and correcting abnormal values of small hydropower generation data according to claim 5, characterized in that: The method for constructing a waveform library includes separating the null values in the small hydropower generation data set, screening out all continuous normal data segments, and sampling using a sliding window technology.
7. The method for identifying and correcting abnormal values of small hydropower generation data according to claim 6, characterized in that: The size of the sliding window is 48, and the sliding step is 1.
8. The method for identifying and correcting abnormal values of small hydropower generation data according to claim 7, characterized in that: The KNN algorithm is used to fill the missing values in the small hydropower generation data set. The method includes calculating the distance between the data point to be filled and each waveform segment in the waveform library, finding k nearest neighboring waveform segments, and filling the missing values according to the weighted average of these neighboring waveform segments.
9. The method for identifying and correcting abnormal values of small hydropower generation data according to claim 8, characterized in that: The KNN algorithm calculates the distance between the sequence to be filled and the normal waveform segments in the waveform library through the Euclidean distance, and calculates the weighting coefficient according to the inverse of the distance. The waveform segments with closer distances obtain greater weights, and finally fill the missing data through the weighted average.
10. A system for identifying and correcting abnormal values of small hydropower generation data, characterized by: include, The data preprocessing module is used to traverse the small hydropower generation data set and replace the zero values with null values; A data change rate calculation module, used to calculate a forward power change rate and a backward power change rate; An outlier identification module is used to set a confidence interval based on the mean and standard deviation of the forward and backward change rates, identify outliers that exceed the confidence interval and perform blanking processing; The data filling module is used to obtain normal data samples from the waveform library and fill in the missing values in the data set.