A cascade power station dispatching operation data cleaning method and system
By establishing a single-item data quality diagnostic model and a multi-item data joint verification model, and combining the water balance principle and characteristic curves, anomalies in hydropower station scheduling and operation data are identified and corrected. This achieves efficient data cleaning and classification, improves data quality and consistency, and supports scheduling decisions for cascade power stations.
Patent Information
- Application Number
- CN202510933943.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-07-08
AI Technical Summary
Inconsistencies and anomalies in hydropower station dispatch and operation data across different time scales and variables lead to difficulties in data processing, affecting the formulation of dispatch plans and the accuracy of power station output plans.
By establishing a single data quality diagnostic model and a multi-data joint verification model, combined with the water balance principle and characteristic curves, data anomalies are identified and corrected. Data is then filtered and classified using classification thresholds to form a high-quality scheduling operation dataset.
It improves the comprehensiveness and accuracy of data processing, ensures that data quality meets actual needs, and supports precise analysis and decision-making for the scheduling and operation of cascade power stations.
Smart Images

Figure CN120430592B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and in particular relates to a method and system for cleaning cascade power station dispatching operation data. Background Art
[0002] Hydropower station operation data is an important basis for developing dispatch plans and analyzing response patterns. However, during the actual operation of reservoirs, random fluctuations and anomalies in the data are inevitable, which can affect the analysis of dispatch element response patterns, the formulation of power station output plans, and the compilation of reservoir dispatch plans.
[0003] During the operation of a power station, flow data is generally on a 2-hour time scale, output data is on a 15-minute time scale, water level data is on a minute-level time scale, and gate opening and closing has no fixed time scale. The mismatch in time scales makes it difficult to predict water levels under complex working conditions.
[0004] Therefore, in theory, data of different time scales and different variables should have inherent connections between them, such as data of different time scales should be consistent, power flow-water level-output data should be consistent, etc. However, in reality, the following problems exist in power station scheduling and operation data:
[0005] (1) The time of data anomaly is different. Abnormal data may appear at different time points. For example, the time of occurrence of the above-mentioned abnormal situations is inconsistent.
[0006] (2) The data anomaly variables are different. There are abnormal data on water level, flow, output, etc. at different sites, with different recording methods and different measurement methods.
[0007] (3) Various data anomalies coexist and are mutually causal. For example, the problem of jagged fluctuations in data will lead to inconsistent data with different calculation calibers, and the problem of missing data will lead to excessive fluctuations in water levels.
[0008] The existence of these data anomaly characteristics makes it difficult to use a unified method to deal with all data anomalies. Therefore, it is necessary to complete data cleaning through perfect single-variable data anomaly diagnosis and correction and multivariate joint verification. Summary of the Invention
[0009] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides a method and system for cleaning the dispatching operation data of a cascade power station.
[0010] The present invention adopts the following technical solution: a method for cleaning cascade power station dispatching operation data, comprising the following steps:
[0011] Acquire historical operation data of the cascade power station, the historical operation data including multiple groups of historical single-item data; establish complex correlations based on the multiple groups of historical single-item data, and obtain multiple groups of historical multiple-item data according to the complex correlations;
[0012] Establishing a single data quality diagnosis model, performing quality analysis on the historical single data to obtain single low-quality data, and cleaning and correcting the data to obtain cleaned historical single data;
[0013] Based on the complex correlation, a multi-data joint verification model is constructed, and the multi-data joint verification model is used to verify the historical multi-data to obtain the verified historical multi-data;
[0014] Integrate the cleaned single-item historical data and the verified multiple-item historical data to obtain complete historical data, establish a data screening condition set, and use the data screening condition set to screen the complete historical data to obtain several scheduling operation data sets;
[0015] The following steps are performed for each scheduling operation data set: based on the data screening condition, a corresponding classification threshold is created, and the scheduling operation data set is classified using the classification threshold to obtain a classified data set.
[0016] In a further embodiment, the multiple groups of historical single data are data information based on spatial scale and time series, and at least include: historical water level data, historical flow data, historical output data and historical gate data.
[0017] In a further embodiment, the process of establishing the complex correlation is as follows:
[0018] Based on the water balance principle, a complex correlation between multiple groups of historical single data is constructed, which is expressed as follows:
[0019] ;
[0020] in, 、 、 、 、 and Respectively represent power stations Corresponding to the changes in the reservoir's inflow, outflow, abandoned water flow, power generation flow, dam front water level and storage capacity, is a function of water level and inflow, For power stations Adjacent power station Inbound traffic, For power stations In the i First acquisition cycle The output of time, For power stations In the second acquisition cycle The output of the first acquisition cycle and the second acquisition cycle The relationship between them is as follows: , i The first acquisition cycle the number of times;
[0021] The complex correlation between multiple groups of historical single data is established based on the characteristic curve, which is expressed as follows:
[0022] ;
[0023] For power stations The downstream water level, For power stations The power generation head water level, For power stations of efforts, For power stations The gate opening, For power stations The gross head of power generation, For power stations The water consumption rate, 、 、 and Both represent curve functions.
[0024] In a further embodiment, the process of establishing the single data quality diagnosis model is as follows:
[0025] Define the historical single data to be tested as , represents the time frame, ,in, For historical water level data, For historical traffic data, For historical output data, is the historical gate data;
[0026] Establish data missing identification model, data anomaly identification model, data rapid change identification model and data abnormal fluctuation model respectively;
[0027] The historical single data to be detected Input into the data missing identification model, data anomaly identification model, data rapid change identification model and data abnormal fluctuation model respectively to determine the historical single data Data points in Is it low-quality data? , ;
[0028] The low-quality data includes at least one of the following situations: missing data, abnormal data, rapid data changes, and abnormal data fluctuations.
[0029] In a further embodiment, the multiple data joint verification models include: a data joint verification model based on a time scale and a data joint verification model based on a physical mechanism.
[0030] In a further embodiment, the time-scale-based data joint verification model is used to verify the consistency of historical single data at different time scales, and the verification process is as follows:
[0031] Define the minimum unit time scale and the maximum unit time scale ,in, , is an integer greater than or equal to 3;
[0032] Get the minimum unit time scale respectively Power Station Historical individual data and the maximum unit time scale Historical individual data ;
[0033] Based on the historical individual data Calculate the maximum unit time scale Power Station Estimated individual data ;
[0034] Combined with the maximum unit time scale Historical individual data , use the following formula to calculate the difference of single data :
[0035] ;
[0036] like , then it means that the corresponding single data is from the smallest unit time scale To the maximum unit time scale There are outliers between is the difference threshold.
[0037] In a further embodiment, the physical mechanism-based data joint verification model is used to verify the relationship between historical flow data, historical water level data, and historical output data of the same time scale. The verification steps are as follows:
[0038] If multiple historical data to be verified have the same elements, a complex correlation relationship based on the water balance principle is used for verification to analyze whether there are any abnormal data.
[0039] If the historical data to be verified are different elements, the complex correlation relationship established based on the characteristic curve is used for verification to analyze whether there is abnormal data.
[0040] In a further embodiment, the data screening condition set includes: screening conditions based on the scheduling operation period , Filter conditions based on gate opening and closing conditions And filtering conditions based on inbound traffic ;
[0041] The scheduling operation data set is expressed in the following form: ,in, Indicates the scheduled operation data set. For filtering conditions, , For historical water level data, For historical traffic data, Provide historical data.
[0042] In a further embodiment, the classification thresholds include: a time threshold, a gate opening / closing state value, and an inflow flow threshold;
[0043] Correspondingly, the classification method of the classification data set is as follows:
[0044] The data set will be scheduled to run based on the time threshold Divide the dispatching and operation data into peak water period, normal water period and dry water period; They are water level data, flow data and output data applicable to the dispatching operation period respectively;
[0045] The running data set will be scheduled based on the open and closed status value of the gate Divided into valve fully closed scheduling operation data, valve fully open scheduling operation data and valve partially open scheduling operation data; They are water level data, flow data and output data applicable to gate opening and closing conditions respectively;
[0046] Schedule the running data set based on the inflow flow threshold Divided into high-flow dispatching and operating data, medium-flow dispatching and operating data, and low-flow dispatching and operating data; They are water level data, flow data and output data applicable to inflow.
[0047] A cascade power station dispatching operation data cleaning system, used to implement the above-mentioned data cleaning method, comprises:
[0048] The first module is configured to obtain historical operation data of the cascade power station, the historical operation data including multiple sets of historical single-item data; establish a complex correlation based on the multiple sets of historical single-item data, and obtain multiple sets of historical multiple-item data according to the complex correlation;
[0049] The second module is configured to establish a single data quality diagnosis model, perform quality analysis on the historical single data to obtain single low-quality data, and clean and correct it to obtain cleaned historical single data;
[0050] The third module is configured to construct a multi-data joint verification model based on complex correlations, and use the multi-data joint verification model to verify multiple historical data to obtain verified historical data.
[0051] The fourth module is configured to integrate the cleaned single-item historical data and the verified multiple-item historical data to obtain improved historical data, establish a data screening condition set, and use the data screening condition set to screen the improved historical data to obtain a plurality of scheduling operation data sets;
[0052] The fifth module is configured to perform the following steps on each scheduling operation data set: based on the data screening conditions, a corresponding classification threshold is created, and the scheduling operation data set is classified using the classification threshold to obtain a classified data set.
[0053] Beneficial effects of the present invention: In view of the uncertainty of data anomalies in time, the present invention obtains the historical operation data of the cascade power station and establishes complex correlations based on multiple groups of historical single data for comprehensive analysis and processing. It can cover abnormal situations occurring at different time points and is not limited to a specific time range, thereby effectively capturing abnormal data that may occur at any time and improving the comprehensiveness of data processing.
[0054] At the same time, considering the diversity of data anomalies in variables, anomalies exist in various data types such as water level, flow, and output at different sites, using different recording and measurement methods. The single data quality diagnosis model in this method can perform quality analysis on various historical single data items (such as historical water level data, historical flow data, historical output data, historical gate data, etc.), accurately identifying low-quality data in different variables. At the same time, the multi-data joint verification model can verify multiple data composed of different variables based on physical mechanisms, comprehensively ensuring the quality of different variable data.
[0055] In response to the complex situation where multiple data anomalies coexist and are causally related to each other, such as data fluctuations, missing measurements, and other situations that influence each other, the method first uses a single data quality diagnosis model to diagnose various low-quality situations in single data, such as missing data, abnormal fluctuations, etc., and then uses a multi-data joint verification model to verify the relationship between multiple data from different dimensions (time scale, physical mechanism, etc.). Through layer-by-layer analysis and processing, the adverse causal relationship between the abnormal data is cut off, making the cleaned data more realistic and accurate.
[0056] A single data quality diagnosis model is established to conduct quality analysis on historical single data from multiple angles such as data missing, anomalies, rapid changes, and abnormal fluctuations, to obtain single low-quality data and clean and correct it. Subsequently, a multiple data joint verification model is constructed based on complex correlations to verify multiple historical data from different time scales (data joint verification model based on time scales) and physical mechanisms (data joint verification model based on physical mechanisms). The synergistic effect of multiple models and multiple steps can more accurately discover and correct problems in the data, and improve the accuracy of data cleaning and correction.
[0057] A data screening condition set is set up, and the complete historical data is filtered according to conditions such as the scheduling operation period, gate opening and closing conditions, and inflow flow to obtain a scheduling operation data set. The scheduling operation data set is further classified using classification thresholds to subdivide various classified data sets such as peak water period, normal water period, dry season, different valve opening and closing states, and different flow levels. This refined classification and screening operation can further optimize the data according to actual operating characteristics and logic, remove data that does not meet the requirements, and make the final data higher in quality, which is more conducive to the subsequent scheduling and operation of cascade power stations and other related work. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 This is the data missing identification model identification diagram of the Huangshan water level in Example 1.
[0059] Figure 2 This is a diagram for identifying rapid changes in the water level of Huangshan in Example 1.
[0060] Figure 3 This is a diagram showing the difference in water level data at different time scales in Example 1.
[0061] Figure 4 is the difference in water level data at different time scales after correction in Example 1.
[0062] Figure 5 1 is a comparison chart of the calculated tail water level and the measured tail water level in Example 1.
[0063] Figure 6 1 is a flow-water level relationship diagram for different scheduling operation periods of Example 1.
[0064] Figure 7 This is a flow-water level relationship diagram based on the gate in Example 1.
[0065] Figure 8 1 is a flow-water level relationship diagram for different flow levels of Example 1. DETAILED DESCRIPTION
[0066] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0067] Example 1
[0068] Carry out data standardization work such as validity analysis, abnormal data diagnosis, data verification, data interpolation and repair, and data smoothing for massive operational data; establish corresponding models and algorithms for massive data diagnosis, cleaning, verification, and interpolation; conduct multiple data joint correlation verifications, and on this basis, try to correct the data and calculation methods according to the joint analysis and verification results to improve data quality.
[0069] This embodiment discloses a method for cleaning cascade power station dispatching operation data, comprising the following steps:
[0070] Acquire historical operation data of the cascade power station, the historical operation data including multiple groups of historical single-item data; establish complex correlations based on the multiple groups of historical single-item data, and obtain multiple groups of historical multiple-item data according to the complex correlations;
[0071] Establishing a single data quality diagnosis model, performing quality analysis on the historical single data to obtain single low-quality data, and cleaning and correcting the data to obtain cleaned historical single data;
[0072] Based on the complex correlation, a multi-data joint verification model is constructed, and the multi-data joint verification model is used to verify the historical multi-data to obtain the verified historical multi-data;
[0073] Integrate the cleaned single-item historical data and the verified multiple-item historical data to obtain complete historical data, establish a data screening condition set, and use the data screening condition set to screen the complete historical data to obtain several scheduling operation data sets;
[0074] The following steps are performed for each scheduling operation data set: based on the data screening condition, a corresponding classification threshold is created, and the scheduling operation data set is classified using the classification threshold to obtain a classified data set.
[0075] In a further embodiment, the multiple sets of historical individual data are data information based on spatial scale and time series, and include at least historical water level data, historical flow data, historical output data, and historical gate data. It can be understood that the spatial scale is the geographic coordinates of the power station and the spatial relationship between adjacent power stations, while the time series can be years, months, days, etc.
[0076] Furthermore, historical water level data, historical flow data, historical output data, and historical gate data are prone to problems such as missing, anomalies, rapid changes, and abnormal fluctuations. Data points need to be located and corrected for these data. Therefore, this embodiment provides a single data quality diagnostic model for accurately locating low-quality data to facilitate cleaning and correction.
[0077] The establishment process of the single data quality diagnosis model is as follows:
[0078] Define the historical single data to be tested as , represents the time frame, ,in, For historical water level data, For historical traffic data, For historical output data, is the historical gate data;
[0079] Establish data missing identification model, data anomaly identification model, data rapid change identification model and data abnormal fluctuation model respectively;
[0080] The historical single data to be detected Input into the data missing identification model, data anomaly identification model, data rapid change identification model and data abnormal fluctuation model respectively to determine the historical single data Data points in Is it low-quality data? , ;
[0081] The low-quality data includes at least one of the following situations: missing data, abnormal data, rapid data changes, and abnormal data fluctuations.
[0082] Furthermore, the analysis process of the data missing identification model is as follows: the data missing identification model is used to traverse the historical single data , determine the data point Is there any missing data? or , then it means the data point missing, Indicates data tag. Combined Figure 1There were some missing values in the water level of Huangshan on March 22, 2015. The recorded values were -99 and nan, which clearly showed that the data were missing.
[0083] Correspondingly, the identification steps of the data anomaly identification model are as follows: For data points and data points , , calculated using the following formula k- Reachable distance :
[0084] ;in, is a data point and data points The Euclidean distance between For data points To its k proximity;
[0085] Then, the data point The local reachable density Expressed as:
[0086] Where, is a data point of k A set of adjacent data points, is the number of data points in the set.
[0087] Corresponding data points The local outlier factor It can be calculated by the following formula:
[0088] ;
[0089] like , then the corresponding data point is an abnormal data point, where is the local outlier threshold.
[0090] By calculating the local reachable density and local outlier factor of data points, we can better capture the local characteristics of the data and identify outliers that differ significantly from surrounding data points within a local range. This addresses the problem that existing technologies are easily affected by the overall distribution characteristics of the data and may not be sensitive enough to local anomalies.
[0091] In a further embodiment, the method for identifying the data changing too fast identification model is as follows:
[0092] The historical single data to be detected is defined as , use the following formula to calculate the rate of change of two adjacent data points of the same type of data on the time frame :
[0093] ,like Then determine the data point The change rate between time frame m and time frame m+1 is too fast. is the time interval between two adjacent data points, is the change rate threshold. For practical applications, please refer to Figure 2 The water level at Phoenix Mountain dropped rapidly by 1m within 1 hour, and then rose by 1m within 1 hour, which is an abnormal situation.
[0094] By calculating the rate of change of adjacent data points and comparing it with the rate of change threshold, rapid changes in data in a short period of time can be discovered in a timely manner, which is of great significance for monitoring the stability of the system and predicting potential failures.
[0095] Finally, the working process of the data abnormal fluctuation model is as follows:
[0096] The calculation of the historical single data to be tested is defined as Standard deviation :
[0097] ,
[0098] for Any data point in , represents the mean of the data points in the window M.
[0099] like , then it means the data point For abnormal fluctuations, is the fluctuation lower limit threshold, is the upper limit of fluctuation.
[0100] Using a moving standard deviation to detect unusual data fluctuations takes into account data fluctuations within a specific time window. Compared to traditional fixed-window standard deviation calculation methods, a moving standard deviation can more promptly reflect the changing trends of data fluctuations and is more sensitive to detecting unusual data fluctuations.
[0101] The defects corresponding to the data points are determined based on the above recognition model, and cleaning is performed on the identified defects, such as data interpolation, data correction, data removal, etc. Correspondingly, the cleaned data correction can adopt weighted averaging, statistical distribution, data fitting, etc.
[0102] In practical applications, the distribution of data may change over time or due to other factors. Detection methods based on local information are better able to adapt to such changes because they only focus on the local neighborhood around the data point rather than relying on global statistical information.
[0103] Therefore, in order to achieve data consistency within the power plant and between adjacent power plants, that is, to correct data quality from multiple dimensions, the complex correlation relationship establishment process of this embodiment includes: a complex correlation relationship constructed based on the water balance principle, which is expressed as follows:
[0104] ;
[0105] in, 、 、 、 、 and Respectively represent power stations Corresponding to the changes in the reservoir's inflow, outflow, abandoned water flow, power generation flow, dam front water level and storage capacity, is a function of water level and inflow, For power stations Adjacent power station Inbound traffic, For power stations In the i First acquisition cycle The output of time, For power stations In the second acquisition cycle The output of the first acquisition cycle and the second acquisition cycle The relationship between them is as follows: , i The first acquisition cycle the number of times;
[0106] The complex correlation relationship established based on the characteristic curve is expressed as follows:
[0107] ;
[0108] For power stations The downstream water level, For power stations The power generation head water level, For power stations of efforts, For power stations The gate opening, For power stations The gross head of power generation, For power stations The water consumption rate, 、 、 and Both represent functions.
[0109] Complex correlations are established from two dimensions: the water balance principle and the characteristic curve, allowing for multi-dimensional data considerations and constraints. On the one hand, the water balance principle links key flow rates, water levels, output data, and other key areas, ensuring the physical rationality of the data. On the other hand, the relationships established with the characteristic curves align with the plant's past operational characteristics, identifying and correcting potential data anomalies from various perspectives. This improves data quality across the board, ensuring that the data truly reflects the plant's operating status and adapts to the impact of various internal and external factors.
[0110] Correspondingly, multiple data joint verification models include: a data joint verification model based on time scale and a data joint verification model based on physical mechanism.
[0111] The time-scale-based data joint verification model is used to verify the consistency of historical single data at different time scales. The verification process is as follows:
[0112] Define the minimum unit time scale and the maximum unit time scale ,in, , is an integer greater than or equal to 3;
[0113] Get the minimum unit time scale respectively Power Station Historical individual data and the maximum unit time scale Historical individual data ;
[0114] Based on the historical individual data Calculate the maximum unit time scale Power Station Estimated individual data ;
[0115] Combined with the maximum unit time scale Historical individual data , use the following formula to calculate the difference of single data :
[0116] ;
[0117] like , then it means that the corresponding single data is from the smallest unit time scale To the maximum unit time scale There are outliers between is the difference threshold.
[0118] For ease of understanding, historical single data Taking water level status data as an example, we can obtain the minimum unit time scale Power Station Water level status data and the maximum unit time scale Water level status data ;
[0119] Using weighted moving average model, linear regression model and machine learning model and other calculation models, based on water level status data Calculate the maximum unit time scale Power Station Water level estimates .
[0120] Furthermore, the water level data difference is calculated using the following formulas: :
[0121] ;like , then it means from the smallest unit time scale To the maximum unit time scale There are water level outliers in the water level data between is the water level difference threshold.
[0122] Combined with practical applications, we further provide examples. The difference distribution between the water level data of different stations at the 1-hour scale data and the average and hourly values of the water level calculated based on the 5-minute scale data is shown in the figure below. Figure 3 As shown in the figure, the difference between water level data at different time scales is mostly around 0, indicating that the water level data at different time scales are relatively consistent. However, for some water levels, there are obvious outliers. For example, the water level difference at different time scales in Wushan, Fengjie, and Wulong all have obvious outliers. Therefore, a joint check is performed on the data whose water level difference at different time scales exceeds the threshold, and the rationality of the 5-minute scale data and the 1-hour scale data is analyzed to correct the data. The corrected output difference is as follows: Figure 4 As shown in the figure, the difference of the corrected data is within a reasonable range, and the obvious abnormal data has been basically corrected.
[0123] In another embodiment, a data joint verification model based on physical mechanisms is used to verify the relationship between historical flow data, historical water level data, and historical output data of the same time scale. The verification steps are as follows:
[0124] If multiple historical data to be verified are the same elements, the complex correlation between historical water level data, historical flow data, historical output data and historical gate data based on the water balance principle is used for verification to analyze whether there is abnormal data; in other words, its verification model is that the complex correlation between historical water level data, historical flow data, historical output data and historical gate data based on the water balance principle is consistent with each other.
[0125] If the historical data to be verified are different elements, the complex correlation relationship established based on the characteristic curve is used for verification to analyze whether there is abnormal data.
[0126] For example, based on the outflow of the reservoir, the tailwater level can be calculated through the tailwater level curve, and the measured tailwater level can be compared with the calculated tailwater level. Figure 5 As shown in the figure, most of the data is concentrated near the 1:1 line, but some scattered points deviate slightly from the 1:1 line. Using the same anomaly diagnosis method, the main abnormal data can be identified and corrected.
[0127] Based on the above single-item data cleaning and correction, and the combined verification of multiple data items, the existing output, flow, and water level data can be constructed into a complete dataset. Based on data classification thresholds, data classification tools can be used to classify data, construct classified datasets, and analyze data patterns under different classification scenarios. Due to the large number of data classification elements and thresholds, the following examples illustrate the results of dataset construction using three scenarios: classification based on scheduling operation period, data classification based on gate opening and closing conditions, and data classification based on inflow flow levels.
[0128] Furthermore, the data screening condition set includes: screening conditions based on the scheduling operation period , Filter conditions based on gate opening and closing conditions And filtering conditions based on inbound traffic ;
[0129] The scheduling operation data set is expressed in the following form: ,in, Indicates the scheduled operation data set. For filtering conditions, , For historical water level data, For historical traffic data, Provide historical data.
[0130] Classification thresholds include: time threshold, gate opening and closing status value and inflow flow threshold;
[0131] Correspondingly, the classification method of the classification data set is as follows:
[0132] The data set will be scheduled to run based on the time threshold Divide the dispatching and operation data into peak water period, normal water period and dry water period; These are the water level data, flow data, and output data applicable to the dispatching operation period. For example, if you set the time thresholds to January 1, June 10, September 10, and December 31, you can define January 1 to June 10 as the peak water period, June 10 to September 10 as the normal water period, and September 10 to December 31 as the dry season.
[0133] The flow-output relationship in different dispatching operation periods is as follows: Figure 6 As shown in the figure, the output is smaller under the same flow conditions, but due to the sufficient water flow during the flood season, the power generation flow is larger and the overall output is larger.
[0134] The running data set will be scheduled based on the open and closed status value of the gate Divided into valve fully closed scheduling operation data, valve fully open scheduling operation data and valve partially open scheduling operation data; They are water level data, flow data and output data applicable to gate opening and closing conditions respectively.
[0135] Further, from Figure 7 It can be seen that the flow-output relationship when the gate is open is quite different from that in other situations. This is because when the gate is open, water is abandoned and the outflow cannot be fully used for power generation. Therefore, its flow-output relationship is different from that under normal power generation.
[0136] Schedule the running data set based on the inflow flow threshold Divided into high-flow dispatching and operating data, medium-flow dispatching and operating data, and low-flow dispatching and operating data; They are water level data, flow data and output data applicable to inflow.
[0137] Furthermore, based on the Three Gorges Reservoir inflow, the data can be divided into classification data sets under different flow levels. Taking the flow threshold of 18,000 m³ / s as an example for data classification, the relationship between dynamic reservoir capacity and dam front water level under different inflow conditions is as follows: Figure 8As shown in the figure. Red points in the figure represent flows greater than 18,000 m³ / s, while blue points represent flows less than 18,000 m³ / s. The figure uses the difference between the Cuntan water level and the Fenghuangshan water level in front of the dam to represent the dynamic reservoir capacity. As can be seen from the figure, when the water level in front of the dam is high and the flow rate is low, the difference between the Cuntan and Fenghuangshan water levels is small. The relationship between the difference between the Cuntan and Fenghuangshan water levels and the Fenghuangshan water level varies significantly under different flow levels. When the Fenghuangshan water level is between 155m and 160m, the relationship between the water level difference and the dam water level varies significantly under different flow levels.
[0138] Example 2
[0139] This embodiment discloses a cascade power station dispatching operation data cleaning system, which is used to implement the data cleaning method described in Example 1, including:
[0140] The first module is configured to obtain historical operation data of the cascade power station, the historical operation data including multiple sets of historical single-item data; establish a complex correlation based on the multiple sets of historical single-item data, and obtain multiple sets of historical multiple-item data according to the complex correlation;
[0141] The second module is configured to establish a single data quality diagnosis model, perform quality analysis on the historical single data to obtain single low-quality data, and clean and correct it to obtain cleaned historical single data;
[0142] The third module is configured to construct a multi-data joint verification model based on complex correlations, and use the multi-data joint verification model to verify multiple historical data to obtain verified historical data.
[0143] The fourth module is configured to integrate the cleaned single-item historical data and the verified multiple-item historical data to obtain improved historical data, establish a data screening condition set, and use the data screening condition set to screen the improved historical data to obtain a plurality of scheduling operation data sets;
[0144] The fifth module is configured to perform the following steps on each scheduling operation data set: based on the data screening conditions, a corresponding classification threshold is created, and the scheduling operation data set is classified using the classification threshold to obtain a classified data set.
Claims
1. A method for cleaning cascade power station dispatching operation data, characterized in that: The following steps are involved: Acquire historical operation data of the cascade power station, the historical operation data including multiple groups of historical single-item data; establish complex correlations based on the multiple groups of historical single-item data, and obtain multiple groups of historical multiple-item data according to the complex correlations; Establishing a single data quality diagnosis model, performing quality analysis on the historical single data to obtain single low-quality data, and cleaning and correcting the data to obtain cleaned historical single data; Based on the complex correlation, a multi-data joint verification model is constructed, and the multi-data joint verification model is used to verify the historical multi-data to obtain the verified historical multi-data; Integrate the cleaned single-item historical data and the verified multiple-item historical data to obtain complete historical data, establish a data screening condition set, and use the data screening condition set to screen the complete historical data to obtain several scheduling operation data sets; The following steps are performed for each scheduling operation data set: based on the data screening conditions, a corresponding classification threshold is created, and the scheduling operation data set is classified using the classification threshold to obtain a classified data set; The process of establishing the complex correlation is as follows: Based on the water balance principle, a complex correlation between multiple groups of historical single data is constructed, which is expressed as follows: ; in, 、 、 、 、 and Respectively represent power stations Corresponding to the changes in the reservoir's inflow, outflow, abandoned water flow, power generation flow, dam front water level and storage capacity, is a function of water level and inflow, For power stations Adjacent power station Inbound traffic, For power stations In the i First acquisition cycle The output of time, For power stations In the second acquisition cycle The output of the first acquisition cycle and the second acquisition cycle The relationship between them is as follows: , i The first acquisition cycle the number of times; The complex correlation between multiple groups of historical single data is established based on the characteristic curve, which is expressed as follows: ; For power stations The downstream water level, For power stations The power generation head water level, For power stations of efforts, For power stations The gate opening, For power stations The gross head of power generation, For power stations The water consumption rate, 、 、 and Both represent curve functions; The multiple data joint verification models include: a data joint verification model based on time scale and a data joint verification model based on physical mechanism; The time-scale-based data joint verification model is used to verify the consistency of historical single data at different time scales. The verification process is as follows: Define the minimum unit time scale and the maximum unit time scale ,in, , is an integer greater than or equal to 3; Get the minimum unit time scale respectively Power Station Historical individual data and the maximum unit time scale Historical individual data ; Based on the historical individual data Calculate the maximum unit time scale Power Station Estimated individual data ; Combined with the maximum unit time scale Historical individual data , use the following formula to calculate the difference of single data : ; like , then it means that the corresponding single data is from the smallest unit time scale To the maximum unit time scale There are outliers between is the difference threshold.
2. A method for cleaning cascade power station dispatching operation data according to claim 1, characterized in that: The multiple groups of historical single data are data information based on spatial scale and time series, and at least include: historical water level data, historical flow data, historical output data and historical gate data.
3. A method for cleaning cascade power station dispatching operation data according to claim 1, characterized in that: The establishment process of the single data quality diagnosis model is as follows: Establish data missing identification model, data anomaly identification model, data rapid change identification model and data abnormal fluctuation model respectively; Input the historical single item data to be tested into the data missing identification model, data anomaly identification model, data rapid change identification model and data abnormal fluctuation model respectively to determine whether the data point in the historical single item data is low-quality data; The low-quality data includes at least one of the following situations: missing data, abnormal data, rapid data changes, and abnormal data fluctuations.
4. A method for cleaning cascade power station dispatching operation data according to claim 1, characterized in that: The data joint verification model based on physical mechanism is used to verify the relationship between historical flow data, historical water level data and historical output data of the same time scale. The verification steps are as follows: If multiple historical data to be verified have the same elements, a complex correlation relationship based on the water balance principle is used for verification to analyze whether there are any abnormal data. If the historical data to be verified are different elements, the complex correlation relationship established based on the characteristic curve is used for verification to analyze whether there is abnormal data.
5. The method for cleaning cascade power station dispatching operation data according to claim 1, characterized in that: The data screening condition set includes: screening conditions based on the scheduling operation period , Filter conditions based on gate opening and closing conditions And filtering conditions based on inbound traffic ; The scheduling operation data set is expressed in the following form: ,in, Indicates the scheduled operation data set. For filtering conditions, , For historical water level data, For historical traffic data, Provide historical data.
6. A method for cleaning cascade power station dispatching operation data according to claim 5, characterized in that: The classification thresholds include: time threshold, gate opening and closing state value and inflow flow threshold; Correspondingly, the classification method of the classification data set is as follows: The data set will be scheduled to run based on the time threshold Divide the dispatching and operation data into peak water period, normal water period and dry water period; They are water level data, flow data and output data applicable to the dispatching operation period respectively; The running data set will be scheduled based on the open and closed status value of the gate Divided into valve fully closed scheduling operation data, valve fully open scheduling operation data and valve partially open scheduling operation data; They are water level data, flow data and output data applicable to gate opening and closing conditions respectively; Schedule the running data set based on the inflow flow threshold Divided into high-flow dispatching and operating data, medium-flow dispatching and operating data, and low-flow dispatching and operating data; They are water level data, flow data and output data applicable to inflow.
7. A cascade power station dispatching operation data cleaning system, used to implement the data cleaning method according to any one of claims 1 to 6, characterized in that: include: The first module is configured to obtain historical operation data of the cascade power station, wherein the historical operation data includes multiple groups of historical single item data; Establishing a complex correlation based on the multiple groups of historical single data, and obtaining several groups of historical multiple data according to the complex correlation; The second module is configured to establish a single data quality diagnosis model, perform quality analysis on the historical single data to obtain single low-quality data, and clean and correct it to obtain cleaned historical single data; The third module is configured to construct a multi-data joint verification model based on complex correlations, and use the multi-data joint verification model to verify multiple historical data to obtain verified historical data. The fourth module is configured to integrate the cleaned single-item historical data and the verified multiple-item historical data to obtain improved historical data, establish a data screening condition set, and use the data screening condition set to screen the improved historical data to obtain a plurality of scheduling operation data sets; The fifth module is configured to perform the following steps on each scheduling operation data set: based on the data screening conditions, a corresponding classification threshold is created, and the scheduling operation data set is classified using the classification threshold to obtain a classified data set.
Citation Information
Patent Citations
Cascade hydropower station output prediction method and system
CN119416937A
Method for short-term generation scheduling of cascade hydropower plants coupling cluster analysis and decision tree
US20200090285A1