A method for extracting characteristic indexes of water and electricity big data

By using a big data feature index extraction method for hydropower, and leveraging parallel time series and feature extraction models, the problem of slow operation and maintenance caused by the complexity of hydropower equipment was solved, thereby improving the speed of equipment status assessment and maintenance efficiency.

CN116776113BActive Publication Date: 2026-01-27SHANGHAI JIAOTONG UNIV +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310614046.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-29
Publication Date
2026-01-27
Estimated Expiration
2043-05-29

AI Technical Summary

Technical Problem

Hydropower units have numerous and complex components and complex failure modes, which makes it slow for maintenance personnel to directly analyze raw data, affecting maintenance efficiency.

Method used

A feature index extraction method for hydropower big data is adopted. By using parallel time series and feature extraction models, the statistical correlation of equipment operating parameters is calculated, important data is screened, computational complexity is reduced, and judgment speed is improved.

Benefits of technology

This accelerated the determination of peak-shaving equipment status parameters, improved maintenance efficiency, and reduced the workload of analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116776113B_ABST
    Figure CN116776113B_ABST
Patent Text Reader

Abstract

The application discloses a kind of water and electricity big data feature index extraction method, it is related to peak shaving equipment technical field, including the operating parameter of the acquisition water and electricity generating unit each equipment, and establish the parallel time series of the operating parameter of water and electricity generating unit each equipment;The operating parameter of water and electricity generating unit each equipment is input to the feature extraction model of pre-constructed peak shaving equipment state parameter based on parallel time series;The operating parameter of water and electricity generating unit each equipment is extracted based on the feature of peak shaving equipment state parameter.The statistical quantity of the operating parameter of each equipment is calculated by multidimensional, and the statistical quantity feature correlation of peak shaving equipment state and the operating parameter of water and electricity generating unit each equipment is calculated, and its test importance degree is beneficial to remove the statistical quantity data interference of low importance degree, reduce the complexity of calculation and analysis process, improve model data calculation efficiency, speed up the peak shaving equipment state parameter judgment result speed, improve repair efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of peak-shaving equipment technology, and in particular to a method for extracting feature indicators from hydropower big data. Background Technology

[0002] Peak-shaving equipment, as a crucial component of the power system, plays a vital role in regulating peak electricity demand and ensuring the safe, stable, and economical operation of the power system. While peak-shaving equipment in hydropower units offers advantages such as convenient start-up and shutdown and easy synchronization during grid connection, frequent start-ups and shutdowns, along with complex operating condition transitions, make hydropower units more prone to failure. A key characteristic of hydropower units is their numerous and complex components, resulting in intricate failure modes and mechanisms with multiple coupled factors, making conventional fault analysis based on circuit or mechanical models quite challenging.

[0003] Currently, assessing the status of peak-shaving equipment requires collecting parameters from various components of the hydropower unit. Evaluation results are obtained by comparing actual vibration with predicted vibration under normal operating conditions. However, hydropower units in hydropower stations have numerous and complex components, and their failure modes are characterized by complex mechanisms, multiple related modules, time delays, and varying conditions. The large number of parameters collected makes it slow for maintenance personnel to directly analyze all the raw data. Furthermore, a large amount of related data cannot be directly filtered, leading to slow assessment of the peak-shaving equipment's status parameters and impacting maintenance efficiency. Summary of the Invention

[0004] Therefore, the technical problem solved by this invention is that, due to the large number of devices and complex structures in hydropower units, their failure modes are characterized by complex mechanisms, many related modules, delay manifestations, and different situations. There are many parameters collected from the devices, and the process of operation and maintenance personnel directly analyzing all the raw data is slow. In addition, there is a large amount of related data that cannot be directly filtered, which leads to slow judgment results on the status parameters of peak-shaving equipment and affects maintenance efficiency.

[0005] To address the aforementioned technical problems, this invention provides the following technical solution: a method for extracting feature indicators from hydropower big data, comprising collecting operating parameters of each device in a hydropower unit and establishing a parallel time series of operating parameters of each device in the hydropower unit; inputting the operating parameters of each device in the hydropower unit into a pre-constructed feature extraction model of peak-shaving equipment status parameters based on the parallel time series; and extracting the operating parameters of each device in the hydropower unit based on the features of the peak-shaving equipment status parameters.

[0006] As a preferred embodiment of the hydropower big data feature index extraction method described in this invention, the parallel time series includes: dividing the time of the operating parameters of the peak-shaving equipment into unit time periods, with each unit time period having the same time interval; numbering the unit time periods according to their chronological order; collecting the operating parameter status quantities of each device in the hydropower unit within each unit time period; and recording the operating parameter status quantities of each device based on the unit time period sequence number to generate a parallel time series of the operating parameters of each device in the hydropower unit.

[0007] As a preferred embodiment of the hydropower big data feature index extraction method described in this invention, the feature extraction model for the peak-shaving equipment status parameters is calculated as follows:

[0008] ,

[0009] in, Feature extraction model representing the state parameters of peak-shaving equipment This indicates the first of the various equipment in the hydropower unit. Taiwan equipment, express The corresponding number Time series of state variables This indicates the sampling time corresponding to the state variable. This indicates the sampling length of the state variable.

[0010] As a preferred embodiment of the hydropower big data feature index extraction method described in this invention, the feature extraction model for the peak-shaving equipment status parameters is simplified, and its calculation expression is as follows:

[0011] ,

[0012] in, This represents the number of channels being calculated, i.e., for scalar state variables. For vector state variables Equals data dimension.

[0013] As a preferred embodiment of the hydropower big data feature index extraction method described in this invention, the statistical quantities of the operating parameters of each device in the hydropower unit are extracted based on the feature extraction of the peak-shaving equipment status parameters. These statistical quantities of the operating parameters of each device include the skewness of the sequence, the length of the sequence, the kurtosis of the sequence, the quantile of the empirical distribution function of the sequence, and the bucket entropy of the sequence. Their calculation expressions are as follows:

[0014] ,

[0015] in, This represents the state vector under a single channel, meaning the statistical process is performed channel by channel. This indicates the first element in the vector. There are several components, where 'a' represents the sequence number. Indicates the sequence length. Represents a sequence skewness, This indicates the kurtosis of the sequence. This represents the standard deviation of the sequence. This represents the empirical distribution function of the sequence. Represents the empirical distribution function of the sequence. quantiles, This represents the bucket entropy of the sequence. Represents the interval number. This indicates the significance of the correlation of statistical characteristics of the operating parameters of various hydropower units within this sequence interval. Indicates the first The probability corresponding to each interval.

[0016] As a preferred embodiment of the hydropower big data feature index extraction method described in this invention, the statistical correlation between the status of peak-shaving equipment and the operating parameters of each piece of equipment in the hydropower unit is calculated, and their importance is tested. A low value indicates a low correlation in the statistical characteristics of the operating parameters of various equipment in a hydropower unit. A high value indicates a high correlation in the statistical characteristics of the operating parameters of various equipment in the hydropower unit; then based on... The value is used to select statistical quantities for the operating parameters of each piece of equipment in the hydropower unit.

[0017] As a preferred embodiment of the hydropower big data feature index extraction method described in this invention, the following steps are included: calculating the statistical correlation between the status of peak-shaving equipment and the operating parameters of each piece of equipment in the hydropower unit, and specifically testing their importance; firstly, based on the different feature types and sample categories, using the Fischer exact test, KS test, and Kendall's rank test to obtain... The value is used to determine whether the hypothesis is true. The statistic for the operating parameters of each device can be represented as n, and the corresponding hypothesis test result can be represented as... The formula for the Fischer exact test is:

[0018] ,

[0019] The expression for calculating the KS test is:

[0020] ,

[0021] The expression for calculating Kendall's rank is:

[0022] ,

[0023] in, Assume that the description of this feature is unrelated to the prediction of the sample class. Assume that this feature is related to the prediction of the sample class, and the number of states represents... and Their variables represent respectively and , and These represent the state sequence number. Represents the total number of samples. express The number of samples, express The number of samples, express The number of samples, Represents the cumulative distribution function. The numbering in the ordered variable sequence represents variables, Represents the total number of samples. These represent the ordered sequences of independent variables corresponding to different categories of dependent variables. Indicates distance calculation, This represents the total number of samples.

[0024] As a preferred embodiment of the hydropower big data feature index extraction method described in this invention, wherein: for The value sequence is sorted, and a threshold is calculated based on the set value of the error detection rate. The error detection rate is calculated by comparing monitoring data of the same type of equipment under both interference-free and interference-affected conditions. Data exceeding the error threshold is considered erroneous, and the error rate of this statistic is used to obtain the error detection rate value. The statistical measures with values ​​below a threshold are used to extract the operating parameters of each piece of equipment in the hydropower unit, where the threshold is... The calculation expression is:

[0025] ,

[0026] in, Indicates the threshold. Indicates according to The assumed index after sorting the values. This represents the error detection rate setting. This represents the total number of assumptions.

[0027] The beneficial effects of this invention are as follows: By statistically analyzing the operating parameters of each device from multiple dimensions, the correlation between the statistical characteristics of the peak-shaving equipment status and the operating parameters of each device in the hydropower unit is calculated, and the importance of these parameters is tested. This helps to remove interference from statistical data with low importance, reduce the complexity of the calculation and analysis process, improve the efficiency of model data calculation, speed up the judgment of peak-shaving equipment status parameters, and improve maintenance efficiency. Attached Figure Description

[0028] Figure 1 This is a schematic diagram of the basic process of a method for extracting feature indicators from hydropower big data, provided as an embodiment of the present invention.

[0029] Figure 2 This is a schematic diagram of the parallel time series establishment process in a method for extracting feature indicators of hydropower big data according to an embodiment of the present invention.

[0030] Figure 3 This is a schematic diagram of the parallel time series feature extraction process of a hydropower big data feature index extraction method provided in one embodiment of the present invention.

[0031] Figure 4 This is a schematic diagram illustrating the distribution of four evaluation levels across different equipment types in one embodiment of a hydropower big data feature index extraction method provided by an embodiment of the present invention. Detailed Implementation

[0032] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0033] Example 1

[0034] Reference Figures 1-3 As an embodiment of the present invention, a method for extracting feature indicators from hydropower big data is provided, comprising:

[0035] S1: Collect the operating parameters of each piece of equipment in the hydropower unit and establish a parallel time series of the operating parameters of each piece of equipment in the hydropower unit.

[0036] In this embodiment, the various components of the hydropower unit can be divided into an electromagnetic unit, a mechanical unit, a control unit, and a comprehensive unit. The electromagnetic unit contains electromagnetic quantities corresponding to the primary equipment, including the partial discharge of the generator and transformer, the air gap of the generator set, the magnetic field strength, and the core grounding current. The mechanical unit contains the state quantities corresponding to the mechanical parts of the turbine and generator, including the vibration of the frame and support cover, the oscillation of the upper guide and water guide, the vibration of the stator core, the water pressure pulsation of the turbine, the degree of cavitation and erosion, and the water pressure at different locations. The control unit contains the monitoring quantities corresponding to the control system, including the head, unit speed, guide vane opening, terminal voltage, stator current, excitation current, active power, reactive power, system-related control parameters and given and switching information. The comprehensive unit contains unit-related temperature quantities such as the upper guide shaft bearing temperature and stator temperature, as well as auxiliary equipment oil level and pressure information.

[0037] Currently, the parameters of each device in the hydropower unit are collected by sensors, and then the sampled signals are aggregated to the peak-shaving equipment to monitor each device in the hydropower unit. For high sampling rate signals, such as partial discharge, the storage space required for all the original data of the device's operating parameters is too large. Therefore, the peak-shaving equipment system generally only retains data records that exceed the limit value. That is, when a certain state quantity exceeds the limit at a certain moment, the system will record data for a fixed length before and after that moment. In addition, in order to obtain the distribution of normal state quantities under different operating conditions, a portion of normal data can be randomly selected for recording.

[0038] The operation parameters of each piece of equipment in the hydropower unit are collected, and a parallel time series of the operation parameters of each piece of equipment in the hydropower unit is established, including:

[0039] The operating parameters of each hydropower unit are input into a pre-built feature extraction model of the peak-shaving equipment status parameters based on parallel time series.

[0040] Operating parameters of each device in a hydropower unit are extracted based on the feature extraction of the status parameters of the peak-shaving equipment.

[0041] Parallel time series include:

[0042] The operating parameters of the peak-shaving equipment are divided into units of time, and the time interval of each unit of time is the same. Since there are large differences in the sequence scale, i.e. the sampling frequency, of different state quantities, and there are correlations between different equipment and different state quantities, the time interval of each unit of time is made the same in order to reduce the error of the collected data.

[0043] The units of time are sequentially numbered according to their order of time; this is to facilitate the organization and calculation of the units of time.

[0044] The operating parameters and status values ​​of each piece of equipment in the hydropower unit are collected within a unit of time; the operating parameters and status values ​​of each piece of equipment in the hydropower unit need to be collected simultaneously within a unit of time.

[0045] Based on the sequence number of unit time, the operating parameter status quantities of each device are recorded to generate a parallel time series of operating parameters for each device of the hydropower unit. The time units corresponding to the sequence numbers and the corresponding operating parameter status quantities of the devices are organized to form a dataset of operating parameter status quantities for each device of the hydropower unit for a certain time unit, which is the parallel time series of operating parameters for each device of the hydropower unit.

[0046] S2: The operating parameters of each device in the hydropower unit are input into a pre-built feature extraction model for the state parameters of the peak-shaving equipment, based on parallel time series. The calculation expression of the feature extraction model for the state parameters of the peak-shaving equipment is as follows:

[0047] ,

[0048] in, Feature extraction model representing the state parameters of peak-shaving equipment This indicates the first of the various equipment in the hydropower unit. Taiwan equipment, express The corresponding number Time series of state variables This indicates the sampling time corresponding to the state variable. This indicates the sampling length of the state variable.

[0049] The feature extraction model for the state parameters of peak-shaving equipment is simplified, and its calculation expression is as follows:

[0050] ,

[0051] in, This represents the number of channels being calculated, i.e., for scalar state variables. For vector state variables Equals data dimension.

[0052] Specifically, this implementation generates parallel time series of operating parameters for each piece of equipment in the hydropower unit, facilitating maintenance personnel to statistically analyze these parameters. Currently, maintenance personnel typically rely on the operating parameters of each piece of equipment. However, due to the large number of operating parameters, directly analyzing all the raw data is slow and cannot be directly filtered, leading to slow judgment of the status parameters of peak-shaving equipment and affecting maintenance efficiency. By extracting the interrelated data from the features of the status parameters of peak-shaving equipment, the workload of analysis and statistics can be reduced, and equipment operating condition information can be quickly obtained to determine whether equipment needs to be replaced or repaired.

[0053] S3: Extracting operating parameters of each device in a hydropower unit based on the feature of the peak-shaving equipment status parameters.

[0054] Based on the feature extraction of peak-shaving equipment state parameters, statistical measures of the operating parameters of each hydropower unit are obtained. These statistical measures include the skewness, length, kurtosis, quantiles of the empirical distribution function, and bucket entropy of the sequence. Their calculation expressions are as follows:

[0055] ,

[0056] in, This represents the state vector under a single channel, meaning the statistical process is performed channel by channel. This indicates the first element in the vector. There are several components, where 'a' represents the sequence number. Indicates the sequence length. Represents a sequence skewness, This indicates the kurtosis of the sequence. This represents the standard deviation of the sequence. This represents the empirical distribution function of the sequence. Represents the empirical distribution function of the sequence. quantiles, This represents the bucket entropy of the sequence. Represents the interval number. This indicates the significance of the correlation of statistical characteristics of the operating parameters of various hydropower units within this sequence interval. Indicates the first The probability corresponding to each interval.

[0057] In this implementation, the statistical measures of the operating parameters of each device can include maximum value, minimum value, mean, variance, and standard deviation. These values ​​can be calculated using existing mathematical formulas to obtain their corresponding statistical measures. Through multi-dimensional calculation and analysis of the statistical measures of the operating parameters of each device, and by assessing feature correlation and testing their importance, the operating parameters of each device can be extracted. This implementation also includes the analysis of the kurtosis, standard deviation, and empirical distribution function of the sequence. The quantiles and bucket entropy of the sequence are correlated, and their importance is tested.

[0058] Calculate the statistical characteristics of the correlation between the status of peak-shaving equipment and the operating parameters of each piece of equipment in the hydropower unit, and test their importance.

[0059] A low value indicates a low correlation in the statistical characteristics of the operating parameters of various equipment in a hydropower unit. A high value indicates a high correlation between the statistical characteristics of the operating parameters of various equipment in a hydropower unit;

[0060] Then based on The value is used to select statistical quantities for the operating parameters of each piece of equipment in the hydropower unit.

[0061] Specifically, this implementation calculates the correlation between the statistical characteristics of the operating parameters of each device and the operating parameters of each device in the peak-shaving equipment and the hydropower unit by using multi-dimensional statistical measures of the operating parameters of each device. It also tests the importance of these statistical measures, which helps to remove interference from statistical data with low importance, reduce the complexity of the calculation and analysis process, improve the efficiency of model data calculation, speed up the judgment of the peak-shaving equipment status parameters, and improve maintenance efficiency.

[0062] Calculate the statistical correlation between the status of peak-shaving equipment and the operating parameters of each piece of equipment in the hydropower unit, and specifically determine the importance of their testing, including:

[0063] First, based on the different feature types and sample categories, the exact test, KS test, and Kendall's rank test are used to obtain... The value is used to determine whether the hypothesis is true. The statistic for the operating parameters of each device can be represented as n, and the corresponding hypothesis test result can be represented as... The formula for the Fischer exact test is:

[0064] ,

[0065] The expression for calculating the KS test is:

[0066] ,

[0067] The expression for calculating Kendall's rank is:

[0068] ,

[0069] in, Assume that the description of this feature is unrelated to the prediction of the sample class. Assume that this feature is related to the prediction of the sample class, and the number of states represents... and Their variables represent respectively and , and These represent the state sequence number. Represents the total number of samples. express The number of samples, express The number of samples, express The number of samples, Represents the cumulative distribution function. The numbering in the ordered variable sequence represents variables, Represents the total number of samples. These represent the ordered sequences of independent variables corresponding to different categories of dependent variables. Indicates distance calculation, This represents the total number of samples.

[0070] Specifically, this implementation calculates the correlation between the statistical characteristics of the peak-shaving equipment status and the operating parameters of each hydropower unit by using multi-dimensional statistical measures of the operating parameters of each equipment. Furthermore, it can verify the correlation of the statistical characteristics of the extracted operating parameters based on actual samples and test their importance by comparing the monitoring data of the same type of equipment under the same interference state and the interference state in advance.

[0071] Regarding the above The value sequence is sorted, and a threshold is calculated based on the set value of the error detection rate. The error detection rate is calculated by comparing monitoring data of the same type of equipment under both interference-free and interference-affected conditions. Data exceeding the error threshold is considered erroneous, and the error rate of this statistic is used to obtain the error detection rate value. The statistical measures with values ​​below a threshold are used to extract the operating parameters of each piece of equipment in the hydropower unit, where the threshold is... The calculation expression is:

[0072] ,

[0073] in, Indicates the threshold. Indicates according to The assumed index after sorting the values. This represents the error detection rate setting. This represents the total number of assumptions.

[0074] By removing interference from statistical data of low importance, the complexity of the calculation and analysis process is reduced, the efficiency of model data calculation is improved, the speed of judging the status parameters of peak-shaving equipment is accelerated, and the maintenance efficiency is improved.

[0075] Example 2

[0076] Reference Figure 3 and 4This is another embodiment of the present invention. Unlike the first embodiment, this embodiment provides an experimental verification of a method for extracting feature indicators from hydropower big data. To verify and explain the technical effects of the method, this embodiment uses a traditional technical solution to compare and test with the method of the present invention. The experimental results are compared using scientific demonstration methods to verify the real effect of the method.

[0077] The operation data of a hydropower station in a certain area, recorded by the computer monitoring system of a certain hydropower company, was selected for analysis and confirmation.

[0078] For data collected from different devices and sensors, a unified monitoring and diagnostic unit time is required; this method uses a daily unit. Daily data from each device and monitoring unit is collected, including active power, reactive power, water supply pressure and direction, component oscillation, stator temperature, cold air temperature, hot air temperature, upper guide bearing temperature and oil level, thrust bearing temperature and oil level, and other relevant parameters. Based on relevant thresholds and trend analysis, the system first makes a preliminary judgment on the status of each device, which is then further confirmed and verified by maintenance personnel.

[0079] For equipment-level condition assessment, a single case refers to different monitoring data records for a specific device on a single day. The dataset contains 2816 cases over a period of 176 days. Each case is categorized into four levels based on equipment type and operating condition: Normal, Attention, Abnormal, and Critical. The number of cases in each category is 1968, 533, 533, and 75, respectively. The distribution of the four assessment levels across different equipment types is shown below. Figure 4 As shown.

[0080] Table 1: Feature extraction verification record table for four levels in different equipment types.

[0081]

[0082] It is evident that the feature extraction model for normal and alert states does not show significant accuracy in extracting statistical correlations of operating parameters of each hydropower unit. This is because the parameter variation range is small in these two states, and the interference between parameters of different devices is relatively strong. However, as the state level increases, the feature extraction model extracts statistical correlations of operating parameters of each hydropower unit. After testing for importance, it is clear that the accuracy of the feature extraction model gradually increases with abnormal equipment parameters. This is beneficial for removing interference from statistical data with low importance, reducing the complexity of the calculation and analysis process, improving the model's data calculation efficiency, accelerating the judgment of peak-shaving equipment status parameters, and improving maintenance efficiency.

[0083] It should be recognized that embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable storage medium. The method can be implemented using standard programming techniques—including a non-transitory computer-readable storage medium configured with a computer program, wherein such a storage medium causes the computer to operate in a specific and predefined manner—according to the methods and drawings described in the specific embodiments. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. Furthermore, for this purpose, the program can run on a programmed application-specific integrated circuit (ASIC).

[0084] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for extracting feature indicators from hydropower big data, characterized in that, include: Collect the operating parameters of each piece of equipment in the hydropower unit and establish a parallel time series of the operating parameters of each piece of equipment in the hydropower unit; The operating parameters of each hydropower unit are input into a pre-built feature extraction model of the peak-shaving equipment status parameters based on parallel time series. Based on the feature extraction of the status parameters of the peak-shaving equipment, the operating parameters of each equipment in the hydropower unit are extracted; The feature extraction model for the state parameters of the peak-shaving equipment has the following calculation expression: , in, Feature extraction model representing the state parameters of peak-shaving equipment This indicates the first of the various equipment in the hydropower unit. Taiwan equipment, express The corresponding number Time series of state variables This indicates the sampling time corresponding to the state variable. This indicates the sampling length of the state variable; Based on the feature extraction of peak-shaving equipment state parameters, statistical measures of the operating parameters of each hydropower unit are obtained. These statistical measures include the skewness of the sequence, the length of the sequence, the kurtosis of the sequence, the quantiles of the empirical distribution function of the sequence, and the bucket entropy of the sequence. Their calculation expressions are as follows: , in, This represents the state vector under a single channel, meaning the statistical process is performed channel by channel. This indicates the first element in the vector. There are several components, where 'a' represents the sequence number. Indicates the sequence length. Represents a sequence skewness, This indicates the kurtosis of the sequence. This represents the standard deviation of the sequence. This represents the empirical distribution function of the sequence. Represents the empirical distribution function of the sequence. quantiles, This represents the bucket entropy of the sequence. Represents the interval number. This indicates the significance of the correlation of statistical characteristics of the operating parameters of various hydropower units within this sequence interval. Indicates the first The probability corresponding to each interval; Calculate the statistical characteristics of the correlation between the status of peak-shaving equipment and the operating parameters of each piece of equipment in the hydropower unit, and test their importance. A low value indicates a low correlation in the statistical characteristics of the operating parameters of various equipment in a hydropower unit. A high value indicates a high correlation between the statistical characteristics of the operating parameters of various equipment in a hydropower unit; Then based on The value is used to select statistical quantities for the operating parameters of each piece of equipment in the hydropower unit.

2. The method for extracting feature indicators from hydropower big data as described in claim 1, characterized in that, The parallel time series includes: The operating parameters of the peak-shaving equipment are divided into units of time, and the time interval between each unit of time is the same. The time units are sequentially numbered according to their chronological order; The operating parameters and status data of each piece of equipment in the hydropower unit are collected within a unit of time. Based on the sequence number of unit time, the operating parameter status of each device is recorded to generate a parallel time series of operating parameters of each device in the hydropower unit.

3. The method for extracting feature indicators from hydropower big data as described in claim 1, characterized in that: The feature extraction model for the state parameters of peak-shaving equipment is simplified, and its calculation expression is as follows: , in, This represents the number of channels being calculated, i.e., for scalar state variables. For vector state variables Equals data dimension.

4. The method for extracting feature indicators from hydropower big data as described in claim 1, characterized in that: Calculate the statistical correlation between the status of peak-shaving equipment and the operating parameters of each piece of equipment in the hydropower unit, and specifically determine the importance of their testing, including: First, based on the different feature types and sample categories, the exact test, KS test, and Kendall's rank test are used to obtain... The value is used to determine whether the hypothesis is true. The statistic for the operating parameters of each device can be represented as n, and the corresponding hypothesis test result can be represented as... The formula for the Fischer exact test is: , The expression for calculating the KS test is: , The expression for calculating Kendall's rank is: , in, Assume that the description of this feature is unrelated to the prediction of the sample class. Assume that this feature is related to the prediction of the sample class, and the number of states represents... and Their variables represent respectively and , and These represent the state sequence number. Represents the total number of samples. express The number of samples, express The number of samples, express The number of samples, Represents the cumulative distribution function. The numbering in the ordered variable sequence represents variables, Represents the total number of samples. These represent the ordered sequences of independent variables corresponding to different categories of dependent variables. Indicates distance calculation, This represents the total number of samples.

5. The method for extracting feature indicators from hydropower big data as described in claim 4, characterized in that: right The value sequence is sorted, and a threshold is calculated based on the set value of the error detection rate. The error detection rate is calculated by comparing monitoring data of the same type of equipment under both interference-free and interference-affected conditions. Data exceeding the error threshold is considered erroneous, and the error rate of this statistic is used to obtain the error detection rate value. The statistical measures with values ​​below a threshold are used to extract the operating parameters of each piece of equipment in the hydropower unit, where the threshold is... The calculation expression is: , in, Indicates the threshold. Indicates according to The assumed index after sorting the values. This represents the error detection rate setting. This represents the total number of assumptions.

Citation Information

Patent Citations

  • Hybrid neural network fault prediction method and system for high-performance computer

    CN113076239A

  • Hydroelectric generating set state degradation evaluation method and system based on multi-source data fusion

    CN115619287A