An abnormal data cleaning method based on time series matching and bidirectional quartile algorithm

Through the combination of timing matching and bidirectional quartile algorithm, the problem of neglecting timing characteristics in the cleaning of abnormal operation data of wind turbine units is solved, and the effective cleaning of stacked and dispersed abnormal data is realized and the identification of abnormal data in the transition area is improved, thereby improving the accuracy of data cleaning.

CN114968999BActive Publication Date: 2025-08-29CHINA THREE GORGES CORPORATION +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210565276.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-23
Publication Date
2025-08-29
Estimated Expiration
2042-05-23

AI Technical Summary

Technical Problem

The existing abnormal operation data cleaning algorithm for wind turbines ignores the timing characteristics and cannot effectively identify abnormal operation data similar to the spatial distribution characteristics of normal data, especially the accumulated power limit data, which leads to the problems of mis-deletion of normal data and misdeletion of abnormal data.

Method used

The method based on timing matching and bidirectional quartile algorithm is adopted to process stacked and dispersed anomaly data through wind speed-wind power timing matching and bidirectional quartile algorithm, and the abnormal data is cleaned using basic and important trend turning points segmentation, wind speed-wind power timing matching metric, rated power elimination coefficient and vertical and horizontal quartile algorithm.

Benefits of technology

Effectively clean the stacked and dispersed abnormal data, identify the abnormal data in the transition area, solve the problem of difficulty in identifying data categories in traditional methods, and improve the accuracy of data cleaning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114968999B_ABST
    Figure CN114968999B_ABST
Patent Text Reader

Abstract

The present invention discloses an abnormal data cleaning method based on time series matching and a two-way quartile algorithm, which belongs to the field of wind turbine power generation technology. The method comprises the following steps: Step 1: Collecting the measured wind speed and wind power data of the wind turbine; Step 2: Identifying abnormal power limit data of the wind turbine; Specifically comprising: Step 21: Using the basic trend turning point and important trend turning point determination algorithm to segment the wind power time series; Step 22: Screening the power limit segmented time series based on the wind speed-wind power time series matching metric; Step 23: Eliminating the rated power segmented time series based on the rated power elimination coefficient; Step 3: Using the two-way quartile algorithm to clean the dispersed abnormal data of wind power. The present invention can effectively clean both accumulated abnormal data and dispersed abnormal data at the same time, and can effectively identify abnormal operation data in the transition area that has similar spatial distribution characteristics to normal data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of wind power generation, and in particular to an abnormal data cleaning method based on time series matching and a two-way quartile algorithm. Background Art

[0002] Affected by various factors such as wind turbine shutdown, power rationing, communication noise and equipment failure, the collected measured data of wind turbines usually contains a large amount of complex abnormal operation data, which cannot be directly used for wind farm operation and maintenance efficiency evaluation, wind power forecasting, wind turbine power generation capacity assessment and operation status monitoring. It is usually necessary to clean the abnormal operation data of wind turbines first. However, the existing wind turbine abnormal operation data cleaning algorithm is usually based on the spatial distribution characteristics of discrete data, ignoring the time series characteristics of the operation data, and cannot effectively identify abnormal operation data with similar spatial distribution characteristics to normal data; and for a large number of accumulated power rationing data samples, traditional data cleaning methods cannot effectively identify the data category of the transition area, which easily causes the problem of erroneous deletion of normal data and omission of abnormal data. In response to the shortcomings of the prior art, the present invention provides an abnormal data cleaning method based on time series matching and bidirectional quartile algorithm, which is suitable for cleaning accumulated abnormal data and dispersed abnormal data, and can effectively identify abnormal operation data with similar spatial distribution characteristics to normal data in the transition area. Summary of the Invention

[0003] The purpose of the present invention is to propose an abnormal data cleaning method based on time series matching and bidirectional quartile algorithm, which is characterized by comprising the following steps:

[0004] Step 1: Collect the measured wind speed and wind power data of the wind turbine;

[0005] Step 2: Use the wind speed-wind power time series matching algorithm to identify abnormal wind turbine power limit data; specifically, the following steps are involved:

[0006] Step 21: Use the basic trend turning point and important trend turning point determination algorithm to segment the wind power time series;

[0007] Step 22: Filter the power-limited segment time series based on the wind speed-wind power time series matching metric;

[0008] Step 23: Eliminate the rated power segmented time series based on the rated power elimination coefficient;

[0009] Step 3: Use the bi-directional quartile algorithm to clean the abnormal data of wind power dispersion.

[0010] The basic trend turning points in step 21 are defined as follows:

[0011] The basic trend turning points are divided into rising trend turning points and falling trend turning points; suppose the time series of wind power is P= <P1,P2,…,P n >, if the wind power P at time t t Satisfy P t-1 ≥P t <P t+1 or P t-1 >P t ≤P t+1 , then the wind power P t is the turning point of the upward trend; if the wind power P at the t moment t Satisfy P t-1 ≤P t >P t+1 or P t-1 <P t ≥P t+1 , then the wind power P t It is a turning point of the downward trend.

[0012] The important trend turning points in step 21 are defined as follows:

[0013] For the basic turning point sequence of wind power time series An important trend turning point P IT With two consecutive basic trend turning points and In the minimum pattern sequence composed of Whether it is an important trend turning point depends on to P IT and The vertical distance D of the line segment is larger. The more likely it is to be an important trend turning point.

[0014] The size of the vertical distance D is affected by the size of the value itself. The important trend turning points of the wind power time series are determined by combining the relative vertical distance RD and the absolute fluctuation amount. The relative vertical distance RD and the absolute fluctuation amount are calculated as shown in formula (1) and formula (2) respectively:

[0015]

[0016]

[0017] Where RD is The relative vertical distance; t1 is the wind power P IT The corresponding time; t2 is the wind power The corresponding time; t3 is the wind power The corresponding moment; ΔP is the absolute fluctuation of wind power; α is the relative vertical distance threshold; β is the absolute fluctuation threshold.

[0018] The step 22 specifically includes the following sub-steps:

[0019] Step 221: using a data time alignment method to extract wind speed segmented time series corresponding to different wind power segmented time series;

[0020] Step 222: Calculate the length L of different segmented time series; the length L of the segmented time series should satisfy L≥γ, where γ is a time length threshold;

[0021] Step 223: Calculate the wind speed-wind power time series matching degree of different segmented time series, and extract the wind power segmented time series with mismatched time series.

[0022] The step 223 is specifically as follows:

[0023] According to the horizontal accumulation distribution characteristics of wind turbine power limit data in the wind speed-wind power scatter plot, the linear fitting method is used to obtain the linear function of wind speed and wind power in different segmented time series; then, the wind turbine power limit data is extracted according to the slope of the linear function. The closer the slope of the linear function is to 0, the greater the possibility that the data in the segmented time series is power limit data. The linear function and slope threshold are shown in formula (3) and formula (4):

[0024] P=a i V+b i (3)

[0025] a i ≤δ (4)

[0026] Where a i is the slope of the linear function corresponding to the i-th segmented time series; b i is the intercept of the linear function corresponding to the i-th segmented time series; δ is the slope threshold used to determine the wind turbine power limit data.

[0027] The formula for eliminating the rated power segmented time series based on the rated power elimination coefficient in step 23 is:

[0028]

[0029] Where, is the rated power rejection factor, P N is the rated power of the wind turbine.

[0030] The step 3 specifically includes the following sub-steps:

[0031] Step 31: Use the longitudinal quartile algorithm to clean the dispersed abnormal data of the longitudinally distributed wind turbines; divide the wind speed into several wind speed intervals according to the interval interval of 0.25m / s, and calculate the abnormal value limit of the wind power in each wind speed interval. The data outside the limit is abnormal data; wherein, the calculation method of the abnormal value limit of the wind power in the i-th wind speed interval is shown in formula (6):

[0032]

[0033] Where, is the lower limit of the abnormal value of wind power in the i-th wind speed interval; is the upper limit of the abnormal value of wind power in the i-th wind speed interval; is the first quantile of wind power in the i-th wind speed interval; is the third quantile of wind power in the i-th wind speed interval; is the interquartile range of wind power in the i-th wind speed interval,

[0034] Step 32: Use the horizontal quartile algorithm to clean the scattered abnormal data of the horizontally distributed wind turbines; divide the wind power into several wind power intervals according to the interval of 25kW, and calculate the abnormal value limit of the wind speed in each wind power interval. The data outside the limit is abnormal data; wherein, the calculation method of the abnormal value limit of the wind speed in the i-th wind power interval is shown in formula (7):

[0035]

[0036] Where, is the lower limit of the abnormal value of wind speed in the i-th wind power interval; is the upper limit of the abnormal value of wind speed in the i-th wind power interval; is the first quantile of wind speed in the i-th wind power interval; is the third quantile of wind speed in the i-th wind power interval; is the interquartile range of wind power in the i-th wind speed interval,

[0037] The beneficial effects of the present invention are:

[0038] The present invention can effectively clean both accumulated abnormal data and dispersed abnormal data at the same time, and can effectively identify abnormal operating data in the transition area that has similar spatial distribution characteristics to normal data, thus solving the problem that the data categories in the transition area between accumulated power-limiting data and normal data are difficult to effectively identify. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1This is a flowchart of the abnormal data cleaning method based on time series matching and two-way quartile algorithm;

[0040] Figure 2 This is the result diagram of abnormal data cleaning based on the two-way quartile algorithm;

[0041] Figure 3 This is the result of abnormal data cleaning based on time series matching and two-way quartile algorithm. DETAILED DESCRIPTION

[0042] The present invention proposes an abnormal data cleaning method based on time series matching and a bidirectional quartile algorithm. The present invention is further described below with reference to the accompanying drawings and specific embodiments.

[0043] Figure 1 This is a flowchart of the abnormal data cleaning method based on time series matching and bidirectional quartile algorithm; the specific implementation steps are as follows:

[0044] (1) Collect the measured wind speed and wind power data of the wind turbine. This case uses the measured wind speed and wind power data of a 1.5MW wind turbine. The data length is about 3 months, and the data time resolution is 10 minutes.

[0045] (2) Using the wind speed-wind power time series matching algorithm to identify abnormal wind turbine power limit data, the specific steps are as follows:

[0046] 1) The wind power time series is segmented using the basic trend turning point and important trend turning point determination algorithm. The definitions of basic trend turning points and important trend turning points are as follows:

[0047] a. Basic trend turning point: According to the basic change trend of three adjacent time points in the time series, the basic trend turning point is mainly divided into an upward trend turning point and a downward trend turning point. Assume that the time series of wind power is P = <P1,P2,…,P n 〉, if the wind power P at time t t Satisfy P t-1 ≥P t <P t+1 or P t-1 >P t ≤P t+1 , then the wind power P t is the turning point of the upward trend; if the wind power P at the t moment t Satisfy P t-1 ≤P t >P t+1 or P t-1 <P t ≥P t+1 , then the wind power P t It is the turning point of the downward trend;

[0048] b. Important trend turning points: The determination of important trend turning points is to reduce the impact of noise points in wind power time series; for the basic turning point sequence of wind power time series An important trend turning point P IT With two consecutive basic trend turning points and In the minimum pattern sequence composed of Whether it is an important trend turning point depends on to P IT and The vertical distance D of the line segment is larger. The more likely it is an important trend turning point, the more the size of the vertical distance is affected by the size of the value itself. Therefore, the present invention proposes a method of combining the relative vertical distance RD and the absolute fluctuation amount to determine the important trend turning points of the wind power time series. The calculation method of the relative vertical distance and the absolute fluctuation amount is shown in formula (1) and formula (2):

[0049]

[0050]

[0051] Where RD is The relative vertical distance; t1 is the wind power P IT The corresponding time; t2 is the wind power The corresponding time; t3 is the wind power The corresponding moment; ΔP is the absolute fluctuation of wind power; α is the relative vertical distance threshold, the purpose of which is to eliminate the influence of noise points on important trend turning points. The value of this embodiment is 0.18; β is the absolute fluctuation threshold, the purpose of which is to eliminate the influence of noise points near zero on important trend turning points. The value of this embodiment is 0.01.

[0052] 2) Screening of power-limited segmented time series based on wind speed-wind power time series matching metric. The specific steps are as follows:

[0053] a. First, the data time alignment method is used to extract the wind speed segment time series corresponding to different wind power segment time series;

[0054] b. Then, calculate the length L of each segmented time series. Since wind turbine power-limited operation typically lasts for a certain period of time, the length of each segmented time series must satisfy L ≥ γ. The duration threshold γ is affected by the data temporal resolution and the actual power-limited operation state of the wind turbine. It needs to be determined based on the actual conditions of different wind turbines. In this embodiment, the value is 6.

[0055] c. Finally, the wind speed-wind power time series matching degree of different segmented time series is calculated, and the wind power segmented time series with mismatched time series are extracted. According to the horizontal stacking distribution characteristics of the wind turbine power limit data in the wind speed-wind power scatter plot, the present invention proposes a linear fitting method to obtain the linear function of wind speed and wind power in different segmented time series, and extract the wind turbine power limit data according to the slope of the linear function. The closer the slope of the linear function is to 0, the greater the possibility that the data of the segmented time series is power limit data. The linear function calculation formula and the slope threshold are shown in formula (3) and formula (4):

[0056] P=a i V+b i (3)

[0057] a i ≤δ (4)

[0058] Where a i is the slope of the linear function corresponding to the i-th segmented time series; b i is the intercept of the linear function corresponding to the i-th segmented time series; δ is the slope threshold used to determine the wind turbine power limit data, which is determined according to different wind turbines and is usually close to 0.

[0059] 3) Elimination of rated power segmented time series based on rated power elimination coefficient; Since the spatial distribution characteristics of wind turbine rated power are similar to the power limit data, it is easy to cause misidentification and needs to be eliminated from the mismatched segmented time series. The rated power elimination method is shown in formula (5):

[0060]

[0061] Where, The rated power rejection coefficient is based on the description in GBT19960.1-2005 Wind Turbine Generator Part 1: General Technical Conditions: "Under normal working conditions, the deviation between the wind turbine power output and the theoretical value should not exceed 10%". P N is the rated power of the wind turbine.

[0062] (3) Use the bidirectional quartile algorithm to clean the abnormal data of wind power dispersion. The specific steps are as follows:

[0063] 1) First, the longitudinal quartile algorithm is used to clean the dispersed abnormal data of the longitudinally distributed wind turbines. The wind speed is divided into several wind speed intervals according to the interval interval of 0.25m / s (the number of data in the wind speed interval is not less than one thousandth of the total data volume), and the abnormal value limit of the wind power in each wind speed interval is calculated. The data outside the limit is considered abnormal data. Among them, the calculation method of the abnormal value limit of wind power in the i-th wind speed interval is shown in formula (6);

[0064]

[0065] Where, is the lower limit of the abnormal value of wind power in the i-th wind speed interval; is the upper limit of the abnormal value of wind power in the i-th wind speed interval; is the first quantile of wind power in the i-th wind speed interval; is the third quantile of wind power in the i-th wind speed interval; is the interquartile range of wind power in the i-th wind speed interval,

[0066] 2) Then, the horizontal quartile algorithm is used to clean the scattered abnormal data of the horizontally distributed wind turbines; the wind power is divided into several wind power intervals according to the interval interval of 25kW (the number of data in the wind power interval is not less than one thousandth of the total data volume), and the abnormal value limit of the wind speed in each wind power interval is calculated. The data outside the limit is considered abnormal data. Among them, the calculation method of the abnormal value limit of the wind speed in the i-th wind power interval is shown in formula (7);

[0067]

[0068] Where, is the lower limit of the abnormal value of wind speed in the i-th wind power interval; is the upper limit of the abnormal value of wind speed in the i-th wind power interval; is the first quantile of wind speed in the i-th wind power interval; is the third quantile of wind speed in the i-th wind power interval; is the interquartile range of wind power in the i-th wind speed interval, Figure 2 This is the result of abnormal data cleaning based on the two-way quartile algorithm.

[0069] Figure 3This is a graph showing the results of abnormal data cleaning based on time series matching and the bidirectional quartile algorithm. Analysis of specific examples demonstrates that the abnormal data cleaning method based on time series matching and the bidirectional quartile algorithm provided by the present invention can effectively clean both accumulated and dispersed abnormal data, and can effectively identify abnormal operating data within transition regions that have similar spatial distribution characteristics to normal data.

[0070] This embodiment is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A method for cleaning abnormal data based on time series matching and bidirectional quartile algorithm, characterized in that: The following steps are involved: Step 1: Collect the measured wind speed and wind power data of the wind turbine; Step 2: Use the wind speed-wind power time series matching algorithm to identify abnormal wind turbine power limit data; specifically, the following steps are involved: Step 21: Use the basic trend turning point and important trend turning point determination algorithm to segment the wind power time series; Step 22: Filter the power-limited segment time series based on the wind speed-wind power time series matching metric; The step 22 specifically includes the following sub-steps: Step 221: using a data time alignment method to extract wind speed segmented time series corresponding to different wind power segmented time series; Step 222: Calculate the length L of different segmented time series; the length L of the segmented time series should satisfy L≥γ, where γ is a time length threshold; Step 223: Calculate the wind speed-wind power time series matching degree of different segmented time series, and extract the wind power segmented time series with mismatched time series; The step 223 is specifically as follows: According to the horizontal accumulation distribution characteristics of wind turbine power limit data in the wind speed-wind power scatter plot, the linear fitting method is used to obtain the linear function of wind speed and wind power in different segmented time series; then, the wind turbine power limit data is extracted according to the slope of the linear function. The closer the slope of the linear function is to 0, the greater the possibility that the data in the segmented time series is power limit data. The linear function and slope threshold are shown in formula (3) and formula (4): P=a i V+b i (3) a i ≤δ (4) Where a i is the slope of the linear function corresponding to the i-th segmented time series; b i is the intercept of the linear function corresponding to the i-th segmented time series; δ is the slope threshold used to determine the wind turbine power limit data; Step 23: Eliminate the rated power segmented time series based on the rated power elimination coefficient; Step 3: Use the bi-directional quartile algorithm to clean the abnormal data of wind power dispersion; The step 3 specifically includes the following sub-steps: Step 31: Use the longitudinal quartile algorithm to clean the dispersed abnormal data of the longitudinally distributed wind turbines; divide the wind speed into several wind speed intervals according to the interval interval of 0.25m / s, and calculate the abnormal value limit of the wind power in each wind speed interval. The data outside the limit is abnormal data; wherein, the calculation method of the abnormal value limit of the wind power in the i-th wind speed interval is shown in formula (6): Where, is the lower limit of the abnormal value of wind power in the i-th wind speed interval; is the upper limit of the abnormal value of wind power in the i-th wind speed interval; is the first quantile of wind power in the i-th wind speed interval; is the third quantile of wind power in the i-th wind speed interval; is the interquartile range of wind power in the i-th wind speed interval, Step 32: Use the horizontal quartile algorithm to clean the scattered abnormal data of the horizontally distributed wind turbines; divide the wind power into several wind power intervals according to the interval of 25kW, and calculate the abnormal value limit of the wind speed in each wind power interval. The data outside the limit is abnormal data; wherein, the calculation method of the abnormal value limit of the wind speed in the i-th wind power interval is shown in formula (7): Where, is the lower limit of the abnormal value of wind speed in the i-th wind power interval; is the upper limit of the abnormal value of wind speed in the i-th wind power interval; is the first quantile of wind speed in the i-th wind power interval; is the third quantile of wind speed in the i-th wind power interval; is the interquartile range of wind power in the i-th wind speed interval, 2. The abnormal data cleaning method based on time series matching and bidirectional quartile algorithm according to claim 1 is characterized in that: The basic trend turning points in step 21 are defined as follows: The basic trend turning points are divided into rising trend turning points and falling trend turning points; suppose the time series of wind power is P= <P1,P2,…,P n >, if the wind power P at time t t Satisfy P t-1 ≥P t <P t+1 or P t-1 >P t ≤P t+1 , then the wind power P t is the turning point of the upward trend; if the wind power P at time t t Satisfy P t-1 ≤P t >P t+1 or P t-1 <P t ≥P t+1 , then the wind power P t It is a turning point of the downward trend.

3. The abnormal data cleaning method based on time series matching and bidirectional quartile algorithm according to claim 1 is characterized in that: The important trend turning points in step 21 are defined as follows: For the basic trend turning point sequence of wind power time series An important trend turning point P IT With two consecutive basic trend turning points and In the minimum pattern sequence composed of Whether it is an important trend turning point depends on to P IT and The vertical distance D of the line segment is larger. The more likely it is to be an important trend turning point.

4. The abnormal data cleaning method based on time series matching and bidirectional quartile algorithm according to claim 3 is characterized in that: The size of the vertical distance D is affected by the size of the value itself. The important trend turning points of the wind power time series are determined by combining the relative vertical distance RD and the absolute fluctuation amount. The relative vertical distance RD and the absolute fluctuation amount are calculated as shown in formula (1) and formula (2) respectively: Where RD is The relative vertical distance; t1 is the wind power P IT The corresponding time; t2 is the wind power The corresponding time; t3 is the wind power The corresponding moment; ΔP is the absolute fluctuation of wind power; α is the relative vertical distance threshold; β is the absolute fluctuation threshold.

5. The abnormal data cleaning method based on time series matching and bidirectional quartile algorithm according to claim 1 is characterized in that: The formula for eliminating the rated power segmented time series based on the rated power elimination coefficient in step 23 is: Where, is the rated power rejection factor, P N is the rated power of the wind turbine.

Citation Information

Patent Citations

  • Alarm associated variable detection method and system based on relevance

    CN106778053A

  • A data cleaning method based on two-dimensional probability density estimation and a quartile method

    CN109918364A

  • Wind power plant output equivalent aggregation model construction method considering power limiting factor

    CN111950131A