Abnormal data identification method and system for time-day-and-frequency-band characteristic engineering
By employing a time-sharing, day-based, and frequency-band-based feature engineering method, the problems of delayed and missed identification of medium- and long-term abnormal data in the power grid were solved, enabling accurate identification and traceability of power grid monitoring data and improving the stability and security of the power grid.
Patent Information
- Application Number
- CN202511516936.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-23
AI Technical Summary
Existing technologies struggle to accurately identify medium- and long-term abnormal data in the power grid, leading to equipment aging and deviations in dispatching strategies. They also pose risks of identification lag and missed detections. Furthermore, fixed thresholds are difficult to adapt to dynamic changes in power grid indicators, affecting the stability and security of the power grid.
By employing a time-based, day-based, and frequency-band-based feature engineering approach, a frequency band data regression model is established through the processing, decomposition, and reconstruction of multi-source data. This model filters core feature sets, quantifies the fluctuation range of abnormal data, and constructs an abnormal data identification system to achieve accurate identification of power grid monitoring data.
It improves the accuracy of identifying anomalies in power grid monitoring data, provides evaluation criteria for tracing the source and attribution, and supports refined management and control of power grid operation and troubleshooting of equipment failures.
Smart Images

Figure CN120995359A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of power grid anomaly identification, and specifically discloses an abnormal data identification method and system based on time-of-day and frequency band feature engineering. BACKGROUND
[0002] With the upgrading of China's power grid to a multi-element collaborative form of source, network, load and storage, and the continuous increase of new energy proportion, power grid enterprises at all levels have significantly enhanced the demand for fine monitoring and technical control of power grid operation status. Power grid core operation technical indicators (such as various types of power generation, key equipment output, load response data, etc.) as irregular time series are the core technical carriers reflecting the power supply and demand balance status, equipment operation health degree and dispatching execution effectiveness. The fluctuations of the above technical indicators are influenced by multiple technical and environmental factors: both technical factors including regional load demand dynamic changes, natural weather conditions, power grid topology structure characteristics, core equipment operation status, etc.; and operation and control factors involving power grid dispatching strategy execution deviation, new energy output prediction accuracy deviation, etc. Among them, the index fluctuation caused by objective natural condition changes, normal equipment operation and maintenance, and reasonable load fluctuation belongs to the normal technical range allowed by the power grid operation; while the index fluctuation caused by unreasonable factors (such as abnormal low power generation caused by equipment fault concealment, output data distortion caused by sensor failure, and load response anomaly caused by dispatching strategy technical execution deviation) will directly interfere with the real-time power balance calculation of the power grid, increase the safety risks of frequency deviation and voltage out-of-limit, and may also lead to new energy consumption obstruction and key equipment overload, and even threaten regional power supply continuity and destroy the overall operation stability of the power grid in severe cases.
[0003] Current control means for power grid core technical indicator anomalies mainly rely on single indicator fixed threshold constraints. For example, set the absolute value threshold of power generation deviation, the fluctuation amplitude threshold of equipment output, and the delay threshold of load response, etc. However, the changes of power grid core technical indicators have significant dynamics, complexity and multi-correlation. For example, new energy power generation is influenced by multiple technical factors such as short-time weather mutation, prediction accuracy deviation and equipment start-stop state, and equipment output data varies with real-time load demand and line transmission capacity of the power grid, so it is difficult for fixed threshold to adapt to the dynamic adjustment of indicators with technical conditions. Inappropriate threshold setting may not only lead to missed or false judgments of abnormal data, but also ignore the coordinated anomaly among multiple indicators, reducing the accuracy and response efficiency of power grid technical control.
[0004] The existing research on the abnormality of the core technical indicators of the power grid still has obvious technical limitations: first, the research scope is focused on the short-term dramatic fluctuations of a single technical indicator, and the analysis is focused on the short-term numerical deviation amplitude of the indicator, and the technical characteristics of the medium and long-term abnormality are not fully explored. Long-term abnormalities often imply deep problems such as equipment aging and scheduling strategy technical bias, which are prone to cause chain risks due to identification lag; second, the feature analysis of the abnormality of the indicator lacks technical subdivision of the time scale. The core technical indicators of the power grid all show significant periodical fluctuation characteristics: short-term fluctuations, medium-term periodic deviations, and long-term structural deviations. Among them, short-term fluctuations can be quickly identified through real-time monitoring, while medium and long-term abnormalities often have identification lag or missed judgment problems due to slow changes and hidden technical causes; third, the technical correlation analysis of the influencing factors and the indicators lacks pertinence. Different technologies and environmental factors have different effects on different indicators at different time scales.
[0005] Therefore, the present application provides an abnormal data identification method and system based on time-of-day and frequency band feature engineering, which aims to quantify the abnormal fluctuation amplitude outside the normal fluctuation range caused by objective technical conditions for power grid monitoring data (such as various types of power generation, key equipment output, load response data, and electricity prices, etc.) at different time scales; and through the time-of-day and frequency band technical design, the feature correlation between the power grid monitoring data and the internal and external technical influencing factors is constructed to accurately identify whether the power grid monitoring data is abnormal, providing technical support for the fine technical control of power grid operation, equipment fault diagnosis, and scheduling strategy optimization. SUMMARY
[0006] The present application aims to provide an abnormal data identification method and system based on time-of-day and frequency band feature engineering, which solves the problem of accurately identifying whether the power grid monitoring data is abnormal. The specific scheme is as follows: An abnormal data identification method based on time-of-day and frequency band feature engineering, comprising: processing initial multi-source data to obtain multi-dimensional time series data; decomposing and reconstructing the multi-dimensional time series data to obtain frequency band multi-dimensional time series data of multiple frequency bands; for each frequency band multi-dimensional time series data, establishing a frequency band data regression model and a core feature set; based on the core feature set of the to-be-evaluated time period, outputting the monitoring data prediction value of the to-be-evaluated time period through the frequency band data regression model, and determining the abnormal monitoring data based on the monitoring data prediction value and the target evaluation data.
[0007] Further, the processing of the initial multi-source data to obtain multi-dimensional time series data comprises: performing integrity check on the obtained initial multi-source data to obtain an integrity check result; the integrity check is implemented by checking the meta information of each type of data; performing data outlier cleaning on the initial multi-source data with the integrity check result being complete to eliminate abnormal data and obtain non-abnormal multi-source data; the data outlier cleaning comprises: performing outlier detection on multi-dimensional data by using Mahalanobis distance and performing outlier detection on single-dimensional data by using Grubbs criterion cycle; performing data missing filling on the non-abnormal multi-source data to obtain complete multi-source data; the data missing filling comprises: performing data filling on spatial data by using Kriging interpolation method, performing data filling on time data by using linear interpolation method, or performing data filling on geographic data by using inverse distance weighting method; dividing load days based on load data in the complete multi-source data to obtain multi-type load days and load day monitoring indicators; dividing other data in the complete multi-source data according to daily granularity to obtain multi-source daily granularity data, and aligning the multi-source daily granularity data with the multi-type load days and the load day monitoring indicators of the corresponding load days to obtain multi-dimensional time series data of time-divided days.
[0008] Further, the processing of the initial multi-source data to obtain multi-dimensional time series data comprises: performing integrity check on the obtained initial multi-source data to obtain an integrity check result; the integrity check is implemented by checking the meta information of each type of data; performing data outlier cleaning on the initial multi-source data with the integrity check result being complete to eliminate abnormal data and obtain non-abnormal multi-source data; the data outlier cleaning comprises: performing outlier detection on multi-dimensional data by using Mahalanobis distance and performing outlier detection on single-dimensional data by using Grubbs criterion cycle; performing data missing filling on the non-abnormal multi-source data to obtain complete multi-source data; the data missing filling comprises: performing data filling on spatial data by using Kriging interpolation method, performing data filling on time data by using linear interpolation method, or performing data filling on geographic data by using inverse distance weighting method; dividing load days based on load data in the complete multi-source data to obtain multi-type load days and load day monitoring indicators; dividing other data in the complete multi-source data according to daily granularity to obtain multi-source daily granularity data, and aligning the multi-source daily granularity data with the multi-type load days and the load day monitoring indicators of the corresponding load days to obtain multi-dimensional time series data of time-divided days.
[0009] Further, the processing of the initial multi-source data to obtain multi-dimensional time series data comprises: performing integrity check on the obtained initial multi-source data to obtain an integrity check result; the integrity check is implemented by checking the meta information of each type of data; performing data outlier cleaning on the initial multi-source data with the integrity check result being complete to eliminate abnormal data and obtain non-abnormal multi-source data; the data outlier cleaning comprises: performing outlier detection on multi-dimensional data by using Mahalanobis distance and performing outlier detection on single-dimensional data by using Grubbs criterion cycle; performing data missing filling on the non-abnormal multi-source data to obtain complete multi-source data; the data missing filling comprises: performing data filling on spatial data by using Kriging interpolation method, performing data filling on time data by using linear interpolation method, or performing data filling on geographic data by using inverse distance weighting method; dividing load days based on load data in the complete multi-source data to obtain multi-type load days and load day monitoring indicators; dividing other data in the complete multi-source data according to daily granularity to obtain multi-source daily granularity data, and aligning the multi-source daily granularity data with the multi-type load days and the load day monitoring indicators of the corresponding load days to obtain multi-dimensional time series data of time-divided days.
[0010] Further, the dividing the other data in the complete multi-source data by day granularity to obtain multi-source day granularity data comprises: reducing sampling of fine-granularity data by a maximum triangle three-bar dimension reduction algorithm to obtain day granularity data; the fine-granularity data refers to data with a sampling frequency less than one day; and interpolating coarse-granularity data by an inverse distance weighted difference algorithm to obtain day granularity data; the coarse-granularity data refers to data with a sampling frequency greater than one day.
[0011] Further, the establishing the frequency band data regression model and screening the core feature set comprises: extracting features from the frequency band multi-dimensional time series data to obtain an internal feature sequence and a plurality of external feature sequences; the internal feature sequence is used to represent the associated change relationship of the monitoring data in the time dimension; and the external feature sequence is used to represent the change of other data; calculating correlation coefficients of the monitoring data sequence with the internal feature sequence and the plurality of external feature sequences respectively, sorting the plurality of feature sequences in descending order of absolute values of the correlation coefficients to obtain a feature importance sequence; constructing a frequency band data regression model, and screening the plurality of feature sequences based on performance of the frequency band data regression model under a plurality of importance threshold values to obtain an initial core feature set; and performing a redundancy check on the initial core feature set based on correlation coefficients between the plurality of feature sequences to eliminate redundant feature sequences to obtain the core feature set.
[0012] Further, the screening the plurality of feature sequences to obtain the initial core feature set comprises: determining a current core feature set based on a current importance threshold value; dividing the current core feature set into a plurality of training sets and a validation set; training an initial frequency band data regression model through the plurality of training sets to obtain a trained frequency band data regression model; verifying the trained frequency band data regression model through the validation set to obtain model performance; determining a new importance threshold value, and repeating the determination of the performance of the trained frequency band data regression model by taking the new importance threshold value as the current importance threshold value; taking a model corresponding to the best performance as the final frequency band data regression model, taking a corresponding current core feature set as the initial core feature set, and taking the current importance threshold value as the final importance threshold value; and the best performance refers to the minimum mean square error of the model.
[0013] Further, when the abnormal monitoring data is abnormal electricity price data, the method further comprises processing the abnormal electricity price data to determine the behavior causing the abnormality.
[0014] Further, the processing of the abnormal electricity price data to determine the behavior causing the abnormality comprises: determining a category to which the abnormal electricity price data belongs; the category comprises a peak day high frequency abnormality, a peak day low frequency abnormality, a peak day trend abnormality, a valley day high frequency abnormality, a valley day low frequency abnormality, a valley day trend abnormality, a flat day high frequency abnormality, a flat day low frequency abnormality, and a flat day trend abnormality; determining one or more category indicators based on the category to which the abnormal electricity price data belongs; and matching the category indicators to corresponding indicator ranges to determine the behavior causing the abnormality.
[0015] The application further provides a time-sharing day and frequency band feature engineering abnormal data identification system using the time-sharing day and frequency band feature engineering abnormal data identification method.
[0016] The application has the following advantages and beneficial effects: The application builds a mapping relationship between monitoring data and internal and external influencing factors by time-sharing day and frequency band, avoids regarding monitoring data fluctuation caused by normal changes of the power grid system as abnormal data, and quantifies abnormal data fluctuation range caused by abnormal behavior by comparing integrated results of each frequency band with target evaluation data.
[0017] The application builds a feature engineering of peak day, valley day, and flat day influencing monitoring data from short-term high frequency items, periodic low frequency items, and long-term trend items based on data decomposition integration technology.
[0018] The application analyzes possible data abnormal behaviors of peak day, valley day, and flat day in short term, periodic, and long term, builds an evaluation index system of abnormal behavior, and provides a quantifiable evaluation standard for traceability and attribution of abnormal data. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1An exemplary flow chart of an abnormal data identification method of time-sharing day and frequency band feature engineering provided by the present application is shown in the figure. Figure 2 An exemplary module diagram of an abnormal data identification system of time-sharing day and frequency band feature engineering provided by the present application is shown in the figure. DETAILED DESCRIPTION
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in connection with the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.
[0021] An abnormal data identification method of time-sharing day and frequency band feature engineering provided by the present application is shown in the figure. Figure 1 The technical solution includes the following steps: Step 1: The data collection and preprocessing module performs processing on the initial multi-source data to obtain multi-dimensional time series data.
[0022] The initial multi-source data refers to the raw data of the power grid system collected, which can include node monitoring historical data, regional meteorological historical data, node load historical data, unit data, fuel data, carbon data, renewable energy output prediction, power grid node topology data, and regional hydrological data, etc. The monitoring data can include power generation data and electricity price data. The regional meteorological historical data includes temperature, wind speed, and irradiance, etc. The power grid node topology structure includes line capacity and node location, etc.
[0023] The initial multi-source data is preprocessed to obtain multi-dimensional time series data. The preprocessing process includes data integrity verification, data outlier cleaning, data missing filling, and load day division, etc.
[0024] The obtained initial multi-source data is subjected to integrity verification to obtain an integrity verification result; the integrity verification is realized by checking the meta information of each type of data. For example, the data integrity verification checks whether the meta information such as time stamp, data source identifier, data quality flag, etc. of each type of data is complete, whether the data source is correct at the 96-point time of monitoring data (such as power generation and on-grid power data and declaration data and clearing data), and whether the logical relationship between data is correctly matched.
[0025] perform data outlier cleaning on the initial multi-source data with the integrity check result being complete, and remove abnormal data to obtain non-abnormal multi-source data; the data outlier cleaning includes performing outlier detection on multi-dimensional data by using Mahalanobis Distance and performing outlier detection on single-dimensional data by using Grubbs Criterion in a loop. The outlier cleaning is used to identify discrete outliers possibly existing in the original data, so as to propose abnormal data and abnormal behaviors in historical data. The multi-dimensional data can include meteorological data, monitoring data of each node and load data. The single-dimensional data can include fuel data and carbon data. After removing the outliers, the data is processed according to the data missing value.
[0026] The Mahalanobis Distance screening process is as follows: assuming that multi-dimensional sequences such as meteorological data, node monitoring data and node load data can be regarded as a data set , the dimensions of different data sets can be different, but the number of days is consistent and is , the Mahalanobis Distance of each point in each data set is calculated respectively: ; wherein is an m-dimensional data set in s days, is the overall mean vector of the data set, is the inverse matrix of the covariance matrix of the data set.
[0027] Based on the input of each data set (S-day m-dimensional meteorological data vector), the output of the Mahalanobis Distance set of each data set can be obtained, denoted as , wherein is the Mahalanobis Distance of the data in the data set on the s day, if the square of the Mahalanobis Distance of the data on a day satisfies: , wherein is the threshold value of the chi-square distribution when the significance level is , it is indicated that the value of the day is an outlier.
[0028] The Grubbs Criterion screening process is as follows: assuming that single-dimensional sequences such as fuel data and carbon data are , the statistic of Grubbs test is calculated: ; wherein represents the arithmetic mean of the single-dimensional sequence, after the statistic of Grubbs test is calculated, it is compared with the Grubbs critical value , the critical value can be obtained by looking up a table according to the sample size S and the significance level, if , and the point farthest from the mean is considered as an outlier. After repeated verification, all the outliers can be output (only one outlier is output each time).
[0029] The complete multi-source data is obtained by filling in the missing data of the non-outlier multi-source data. The missing data filling includes filling in the data of the spatial data by using the Kriging interpolation method, filling in the data of the time data by using the linear interpolation method, or filling in the data of the geographic data by using the inverse distance weighting method. The spatial data can include meteorological data. The time data can include monitoring data and load data.
[0030] The Kriging interpolation method steps are: inputting each climate sequence and its spatial position coordinates after anomaly screening into the Kriging interpolation method, and outputting the predicted value of the missing and anomaly deleted points and the estimated variance representing the prediction error estimate. The linear interpolation method steps are: identifying the missing points of the node monitoring data, node load, fuel data sequence after anomaly screening and the values of the known points near the missing points, and filling in the missing values according to the relative days between the known points and the missing points according to the linear interpolation method.
[0031] The multi-class load days and load day monitoring indicators are obtained by dividing the load days based on the load data in the complete multi-source data. The load day division refers to dividing the monitoring data into electricity peak days, electricity valley days and electricity flat days according to the load demand data. The time of the present invention refers to the 96-point time of the electricity market, and the day refers to the peak / valley / flat day.
[0032] The steps of peak-valley-flat day division are: firstly, identifying the maximum value and the average value of the load of each day 96 time, and sorting all the maximum load and average load respectively. According to experience, the load ratio threshold is set to determine the load peak-valley-flat day, for example, set to 10%, that is, the load day corresponding to the first 10% of the maximum load or average load is regarded as the load peak day, and the load day corresponding to the last 10% of the maximum load or average load is regarded as the load valley day. Finally, the union of the peak days determined according to the maximum load and the peak days determined according to the average load is taken as the final load peak day, and the load valley day is determined in the same way, and the remaining days are flat days.
[0033] The monitoring data of each day is in the form of summation or weighted average based on different types. Taking the power generation as an example, the time division monitoring data is:
[0034] wherein, represents the power generation sum of the peak / valley / flat day; represents the power generation of the s day time t; t represents the time variable; T represents the total number of time, for example, 96; s represents the date variable, time represents the total number of days monitored.
[0035] Taking electricity price as an example, the time-weighted average monitoring data of each day is: ; Among them, represent the weighted price of peak / valley / flat days, and t represents 96 time points; represents the load demand at time t on day s; the load demand of each day is the sum of the load at each time of the day; represents the electricity price at time t on day s.
[0036] Other data in the complete multi-source data is divided by day granularity to obtain multi-source day granularity data, and the multi-source day granularity data is aligned with the corresponding load day multi-class load day and load day monitoring indicators to obtain time-division day multi-dimensional time series data, including: fine-grained data is reduced by a maximum triangle three-bucket dimension reduction algorithm to obtain day granularity data; fine-grained data refers to data with a sampling frequency less than one day; coarse-grained data is interpolated by an inverse distance weighted difference algorithm to obtain day granularity data; coarse-grained data refers to data with a sampling frequency greater than one day.
[0037] For example, fine-grained meteorological data is down-sampled to the same granularity as the weighted average monitoring data and load data. The present application uses the LTTB (Largest Triangle Three Buckets) dimension reduction algorithm to adjust the fine-grained meteorological data (e.g., sampling frequency of 1 hour) to daily sampling, retaining the details of the meteorological time series data while ensuring consistency of the data in the time scale, meeting the subsequent data analysis and modeling requirements.
[0038] The steps of the LTTB dimension reduction algorithm include: The various types of fine-grained sampling data of each meteorological station are processed by the LTTB algorithm. First, each sequence is divided into "buckets" of approximately equal size in order, with each bucket containing a certain number of consecutive data points. The number of buckets and the number of data points in each bucket are determined by a set parameter. In order to down-sample the meteorological data to each day, this parameter is set to allow each bucket to contain all fine-grained data for each day. The first and last buckets only contain the first and last data points of the original data.
[0039] Next, the LTTB algorithm traverses all the storage buckets and selects a point from each bucket. The first storage bucket contains only one point, so that point will be selected. Then, the algorithm will process three buckets at a time, selecting the representative point from left to right in chronological order: The next bucket of the selected representative point uses the determined representative point, and the next bucket of the selected representative point uses the average value of all points as a temporary representative point, at this time, the area of the triangle formed by each point in the middle bucket (the current bucket) and the front and rear representative points is calculated, the point forming the largest triangular area is selected as the final representative point of the current bucket.
[0040] The LTTB extracts the most representative data in each bucket as the output sequence by the maximum triangle method, that is, the representative value of the daily meteorological data is finally determined, while the characteristics of the original data are maximally retained.
[0041] Considering that the meteorological conditions are different at different geographical positions, at the geographical space level, the present application adopts the inverse distance weighted method (IDW) to interpolate the discrete meteorological station data near the node, inputs the data of the known meteorological stations near the node, combines the power grid topological structure data, obtains the unknown meteorological data of the geographical position of the node, and constructs the corresponding relationship between the geographical position of the power grid node and the meteorological grid network.
[0042] The basic steps of the inverse distance weighted interpolation method include: After the meteorological data is down-sampled, the daily meteorological representative value of each meteorological station is obtained as the known point position and observation value in the inverse distance weighted interpolation. For each node, first, identify the meteorological stations near the node, and the identification logic is the five meteorological stations closest to the node. Input the meteorological data and position information near the node and the position information of the node into the inverse distance weighted interpolation method: ; Wherein, is the estimated meteorological value of the geographical position of the node n, is the known meteorological value of the adjacent meteorological station, and R is the number of adjacent stations. is the Euclidean distance from the rth meteorological station to the node.
[0043] In this way, the meteorological value of the node is estimated by weighted calculation according to the principle that the closer the meteorological station to the node, the greater the influence on the node.
[0044] The feature sequence of the fuel data and the carbon data is day granularity data, and the granularity after the weighted average of the monitoring data is the same, and it is assumed that the influence degree on each node is the same.
[0045] Step two, the data decomposition and reconstruction module is executed, the multi-dimensional time series data is decomposed and reconstructed to obtain frequency band multi-dimensional time series data of multiple frequency bands, including: Each type of load day is subjected to multivariate empirical mode decomposition to obtain the connotation modal component and the trend item of the multi-type load day; the connotation modal component includes multiple frequencies, including: Based on the monitoring data and the correlation data, a multi-dimensional time sequence is constructed to obtain a multi-dimensional time sequence feature. For peak day, valley day and flat day, a multi-dimensional time sequence of monitoring data and its supply and demand influencing factors (load, meteorology, fuel, carbon emission) is constructed. Taking power generation data as an example, the multi-dimensional time sequence can be , wherein, is a power generation sequence, is an influencing factor sequence, and time represents a peak / valley / flat day. Taking electricity price data as an example, the multi-dimensional time sequence can be , wherein is an electricity price sequence, is an influencing factor sequence, and time represents a peak / valley / flat day (since the traceability analysis of electricity price is respectively performed for peak, valley and flat days, the analysis steps of each day are the same, and in order to be simple, the subscript of time is omitted in the following). The multi-dimensional time sequence feature is decomposed by a multi-component empirical mode decomposition (MEMD) algorithm to obtain connotative modal components of different frequencies and a trend item. Lempel-Ziv algorithm reconstruction is performed to obtain a high-frequency item and a low-frequency item, including: the complexity of each type of load day connotative modal component is calculated by Lempel-Ziv algorithm; the connotative modal components of each type of load day are sequentially added according to the complexity, and the sum of the first order number of connotative modal components of each type of load day and the sum of the second order number of connotative modal components are taken as the high-frequency item and the low-frequency item; the first order number is the number of the minimum connotative modal component whose complexity sum is greater than the reconstruction threshold, starting from the connotative modal component with the maximum complexity of each type; and the second order number refers to the number of connotative modal components other than the first order number.
[0046] The Lempel-Ziv algorithm reconstruction is performed on the multiple connotative modal components of each type of load day to obtain a high-frequency item and a low-frequency item of each type of load day. Taking electricity price data as an example, the multi-dimensional time sequence of each type of load day is input into the MEMD method, and the decomposition modules and the trend item of each day are output, wherein is the number of decomposition modules, are the decomposition values of electricity price and influencing factors on the first decomposition module, respectively, are the decomposition values of electricity price and influencing factors on the second decomposition module, respectively, are the decomposition values of electricity price and influencing factors on the third decomposition module, respectively, are the trend items of electricity price and influencing factors, respectively, , and have the same shape, i.e., the same number of rows and columns.
[0047] The Lempel-Ziv algorithm reconstruction steps are as follows: calculate each The lz value of each IMF is used to determine the high-frequency term. When the cumulative lz value of each IMF exceeds a threshold, the sum of the corresponding IMFs is considered the high-frequency term. The threshold is typically 0.8. Specifically: ; If the above formula is satisfied, then to The sum of the values is high frequency, and the sum of the remaining values of the IMF is low frequency. The trend term represents the trend of the sequence. In this invention, the frequency band refers to high frequency, low frequency, and the trend term.
[0048] The high-frequency, low-frequency, and trend items for each type of load day are used as multi-dimensional time-series data for the frequency band.
[0049] Step 3: Execute feature engineering by time period and frequency band. For each frequency band's multi-dimensional time-series data, establish a frequency band data regression model and select core feature sets. For high-frequency items of short-term fluctuations, low-frequency items of medium-term fluctuations, and trend items of long-term trends, construct feature engineering for each time period to determine the regression relationship between the monitoring data of each frequency band and its supply and demand influencing factors. Feature engineering is used to determine the degree to which each decomposition module of the monitoring data is affected by internal and external influencing factors. This is expressed as the ranking of the correlation coefficients between features and monitoring data in each decomposition module. Decomposition modules include peak day decomposition modules, trough day decomposition modules, and flat day decomposition modules; each type of decomposition module includes high-frequency items, low-frequency items, and trend items.
[0050] Feature extraction is performed on the multi-dimensional time-series data of the frequency band to obtain internal feature sequences and multiple external feature sequences. The internal feature sequences represent the correlation and change relationships of the monitoring data in the time dimension, while the external feature sequences represent the changes in other data. For example, time-series features, load features, meteorological features, fuel features, and carbon emission features are extracted respectively. Among them, time-series features are internal influence features of the monitoring data, and the other features are external influence features. For example, as monitoring data is a time series, its internal influence features refer to the lag values of the monitoring data decomposition module itself; specifically, the high-frequency items of the monitoring data on day t may be affected by the high-frequency items of the monitoring data on day t-1. External influence features refer to factors other than the lag values of the monitoring data itself, including weather conditions such as temperature and wind speed, as well as fuel data and carbon data.
[0051] The internal feature sequence refers to the historical lag value of each node monitoring data sequence in each decomposition module term. The lag value threshold of the monitoring data decomposition time sequence is determined based on the ACF value and the PACF value. The historical data of each lag time interval is extracted with the lag value threshold as the upper limit. For example, for the historical lag value of the high-frequency monitoring data, the value of the monitoring data high-frequency term at t day may be affected by t-1 day and t-2 day; and the low-frequency term may be affected by t-1 day to t-10 day. For another example, if the lag value threshold is determined to be 5 according to the PACF and ACF values, the correlation between the t day monitoring data and the monitoring data sequence from t-1 day to t-5 day is constructed, that is, the monitoring data at t day is predicted by the monitoring data sequence from t-1 day to t-5 day.
[0052] The correlation coefficients of the monitoring data sequence and the internal feature sequence and the plurality of external feature sequences are calculated, the plurality of feature sequences are sorted in descending order of the absolute values of the correlation coefficients, and a feature importance sequence is obtained. The monitoring data sequence refers to a sequence composed of the values of the monitoring data. For example, all the extracted features are calculated based on the Pearson correlation coefficient to obtain the importance of each feature in affecting the monitoring data. Then, the features are sorted in descending order of importance to form a feature importance sequence.
[0053] The expression of the Pearson correlation coefficient is as follows: ; Wherein, PCC represents the Pearson correlation coefficient value; and are the monitoring data and a certain feature sequence, and are the average values of the two sequences.
[0054] A frequency band data regression model is constructed, and based on the performance of the frequency band data regression model under the importance threshold of multiple scales, a plurality of feature sequences are screened to obtain an initial core feature set, including: determining a current core feature set based on a current importance threshold; dividing the current core feature set into a plurality of training sets and a validation set; training an initial frequency band data regression model through the plurality of training sets to obtain a trained frequency band data regression model; verifying the trained frequency band data regression model through the validation set to obtain a model performance; determining a new importance threshold, and taking the new importance threshold as the current importance threshold, repeating the determination of the performance of the trained frequency band data regression model; taking the model with the best performance as the final frequency band data regression model, taking the corresponding current core feature set as the initial core feature set, and taking the current importance threshold as the final importance threshold; the best performance refers to the minimum mean square error of the model.
[0055] For example, the K-fold cross-validation method is used to cross-check each frequency band module to determine the optimal threshold value for feature screening. According to the threshold value, the feature screening is performed to determine the influence of each feature on the monitoring data in each time period and each frequency band. The optimal threshold value refers to the threshold value of the correlation between the feature and the monitoring data. For example, if the threshold value is set to 0.5 according to the cross-checking, the prediction model performs best, which indicates that only the features with a correlation greater than 0.5 are selected, and the features with a correlation less than 0.5 are deleted. The correlation coefficient between the monitoring data and the features in each time period and each frequency band is determined. Each time period refers to the high frequency, low frequency, and trend term in the peak, valley, and flat period.
[0056] The specific steps of the K-fold cross-validation method are as follows: The training set of each decomposition module is divided into multiple folds (e.g., 5 folds or 10 folds, etc.). During each validation, one fold is used as the validation set, and the remaining folds are used as the training set to ensure that each fold has the opportunity to be used as the validation set. Next, a threshold range is set. Determine the threshold range to be tried. For example, start from 0.1 and increment by 0.1 to 0.9 to form multiple different threshold points for subsequent evaluation of the impact of different thresholds on model performance. For each determined threshold, use a simple machine learning model (such as LSTM, ELM, XGBoost) to train on the training set and predict on the validation set. The performance of the model at this threshold is evaluated by calculating the prediction error (such as mean square error), and the model performance index corresponding to each threshold is recorded. According to the relationship between the threshold and the model performance analysis result, select the threshold that can make the model perform best (such as the smallest mean square error) on the validation set as the best threshold. According to the determined best threshold, the sorted feature importance sequence is screened, and those features with importance higher than the threshold are selected as core features. For different frequency band decomposition modules, construct a feature matrix for each time period.
[0057] Based on the correlation coefficients between multiple feature sequences, perform redundancy checking on the initial core feature set to eliminate redundant feature sequences and obtain the core feature set. For example, perform redundancy checking on the features based on the Pearson correlation coefficient. Calculate the correlation coefficient between the features. For highly correlated feature pairs, retain one representative feature and eliminate the other. For example, in addition to the lag value of the monitoring data itself, if the correlation between temperature and humidity is strong, select the one with the highest correlation coefficient with the monitoring data and delete the other. Determine the regression model of the monitoring data in each time period and each frequency band using the Pearson correlation coefficient as the parameter.
[0058] Step four, the abnormal data identification module is executed, based on the core feature set of the to-be-evaluated period, the monitoring data prediction value of the to-be-evaluated period is output through the frequency band data regression model, and based on the monitoring data prediction value and the target evaluation data, the abnormal monitoring data is determined. The regression model of step three describes the change of the monitoring data affected by the normal supply and demand fluctuation of the market, and the evaluation logic of the abnormal monitoring data is that the fluctuation of the monitoring data caused by the normal supply and demand change is the abnormal monitoring data.
[0059] For example, the monitoring data prediction value of each frequency band and each period is calculated respectively, and after integration, it is compared with the target monitoring data. Each frequency band and each period refers to high, medium and low frequency bands, and peak, valley and flat periods. An abnormal monitoring data evaluation index is designed, and if the abnormal index exceeds the threshold value, it means that the monitoring data is abnormal. The threshold value refers to the threshold value of the index, which can be determined according to a large amount of historical data, expert experience and actual market conditions. For example, the abnormal index value of the electricity price is calculated to be 1, and the threshold value is 0.5, so the electricity price data is abnormal. The monitoring data abnormality evaluation index is: ; Among them, is the monitoring data abnormality index of each period, and are the integrated monitoring data prediction value and the target evaluation monitoring data (the actual monitoring data monitoring value), and if the monitoring data abnormality index exceeds the threshold value, it is abnormal data.
[0060] When the abnormal monitoring data is abnormal electricity price data, step five is further included, which processes the abnormal electricity price data to determine the behavior causing the abnormality. It includes: determining the category to which the abnormal electricity price data belongs; the categories include peak day high frequency abnormality, peak day low frequency abnormality, peak day trend abnormality, valley day high frequency abnormality, valley day low frequency abnormality, valley day trend abnormality, flat day high frequency abnormality, flat day low frequency abnormality and flat day trend abnormality; based on the category to which the abnormal electricity price data belongs, one or more category indexes are determined; the category index is matched to the corresponding index range to determine the behavior causing the abnormality.
[0061] For example, an index system of abnormal electricity price influencing factors is constructed for each frequency band and each period, and the basic information of the unit and the unit declaration information data are used to identify the abnormal behavior causing the abnormal electricity price. For example, these indexes will have a threshold value determined by experience or data, and if the actual calculation is outside the threshold value, it means that there is such abnormal behavior. For example: if the peak time high frequency item has abnormal electricity price, the evaluation index of the peak time high frequency can be calculated to determine which behavior causes the abnormality, so as to realize abnormality tracing and behavior positioning.
[0062] Peak day high frequency evaluation index. The sequence characteristics corresponding to the peak time high frequency are large load fluctuations, and the sequence presents short-term dramatic fluctuations. The corresponding behavior is the short-term abnormal behavior at the peak time, such as reporting high price at the peak time, sudden shortage of available capacity outside the plan, etc. The short-term abnormal behavior of electricity peak period causing abnormal electricity price includes economic retention, physical retention, and congestion arbitrage. The specific indicators are: Economic retention index: Peak day declared price deviation rate = (peak day reported price - short-term marginal cost) / short-term marginal cost; Physical retention index: Unit capacity sudden drop rate = (rated capacity - peak day actual available capacity) / rated capacity; Congestion arbitrage: Artificial congestion yield rate = (peak day congestion yield - normal congestion yield average) / normal congestion yield average; Among them, the determination of normal congestion yield average is based on the virtual clearing of units in peak day according to the declared cost.
[0063] Peak day low frequency evaluation index. The characteristics corresponding to the peak time low frequency are large load fluctuations, and the sequence presents periodic fluctuations. The corresponding behavior is the periodic abnormal behavior that may occur at the peak time. For example, seasonal capacity retention is achieved through maintenance. The periodic abnormal behavior of electricity peak day causing abnormal electricity price includes seasonal capacity retention, and electricity generating units intentionally delaying unit commissioning or maintenance before the high temperature or cold wave season to create seasonal shortage. The specific indicators are: Peak day maintenance deviation degree = (actual maintenance start time - historical maintenance start time average) / standard deviation; The standard deviation is calculated based on the deviation day sequence of historical maintenance start time and historical maintenance start time average.
[0064] Peak day maintenance extension rate = (actual maintenance duration - declared maintenance duration) / average value of maintenance of the same type unit at the same period; Peak day trend evaluation index. The trend abnormal behavior of electricity peak day causing abnormal electricity price includes artificially reducing the flexibility of resources and actively maintaining the transmission bottleneck. The specific indicators are: Flexibility gap growth rate = Δ (peak day load demand - adjustable resource capacity) / year; ; Among them, is the annual congestion yield growth rate, and the trend of transmission congestion yield is the three-year moving average of the annual congestion yield growth rate.
[0065] Valley day high-frequency evaluation index. The short-term abnormal behavior of abnormal electricity price in valley period of electricity consumption causes the group containing energy storage and renewable energy. The excess predicted output of renewable energy further lowers the price in the valley period. Due to the pressure of consumption, energy storage can arbitrage through charging. The specific index is: Valley day charging arbitrage index = corr (positive error of renewable energy output prediction, battery charging scalar); Wherein, corr represents the correlation coefficient.
[0066] Valley day low-frequency evaluation index. The periodic abnormal behavior of abnormal electricity price in valley day of electricity consumption includes the scheduling control of hydraulic power to affect the electricity price. The specific index is: Hydrological scheduling deviation index = (actual water release - optimized scheduling amount) / basin average; Valley day trend evaluation index. The trend abnormal behavior of abnormal electricity price in valley day of electricity consumption includes the inhibition of negative electricity price conduction, the delay of energy storage economy, and the maintenance of traditional power source income. The specific index is: Negative electricity price inhibition index = 1 - (number of markets allowing negative electricity price market / total number of markets); Wherein, the number of markets allowing negative electricity price market refers to the number of markets allowing market clearing price to be negative.
[0067] High-frequency evaluation index of flat day. The moderate load demand of electricity consumption in flat day causes short-term abnormal behavior of abnormal electricity price, including unit operation climbing ability to cause short-term supply and demand imbalance. The specific index is: Flat day climbing manipulation index = |approved climbing rate - actual climbing rate| / technical limit; Low-frequency evaluation index of flat day. The periodic abnormal behavior of abnormal electricity price in flat day of electricity consumption includes intentionally underreporting climbing ability in spring and autumn transition season of flat day, delaying unit start-stop response speed, and prolonging the operation time of high-price unit. The evaluation index of flat day low frequency is mainly the transition season price difference elasticity: Transition season price difference elasticity = |(maximum value of transition season flat day electricity price - minimum value of transition season flat day electricity price) / load change rate; Wherein, the transition season mainly refers to the spring and autumn seasons.
[0068] Trend evaluation index of flat day. The trend abnormal behavior of abnormal electricity price in flat day of electricity consumption includes the lagging of climbing standard. The specific index is: Standard lag coefficient = (current standard - international advanced standard) / advanced standard; Through the establishment of the index system in step five, the behavior tracing and responsibility positioning of abnormal electricity price in short, medium and long term in peak day, valley day and flat day are carried out respectively. Finally, a complete market force tracing path is formed, the behavior attribution and mechanism identification are clear, and the empirical basis is provided for post-supervision and policy optimization.
[0069] The application also provides a time-of-day and frequency band feature engineering-based abnormal data identification system, which comprises a data collection and preprocessing module, a data decomposition and reconstruction module, a time-of-day and frequency band feature engineering construction module and an abnormal data identification module. Figure 2 The data collection and preprocessing module is used for processing initial multi-source data to obtain multi-dimensional time series data; the data decomposition and reconstruction module is used for decomposing and reconstructing the multi-dimensional time series data to obtain frequency band multi-dimensional time series data of multiple frequency bands; the time-of-day and frequency band feature engineering construction module is used for establishing a frequency band data regression model and screening a core feature set for each frequency band multi-dimensional time series data; and the abnormal data identification module is used for outputting a monitoring data prediction value of an evaluation period to be evaluated through the frequency band data regression model based on the core feature set of the evaluation period to be evaluated, and determining abnormal monitoring data based on the monitoring data prediction value and target evaluation data.
[0070] The above merely describes the preferred embodiments of the present application, but should not be used to limit the present application. Various modifications and changes can be made by those skilled in the art based on the principles and spirit of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A time-sharing day frequency band feature engineering abnormal data identification method, characterized in that, The method comprises the following steps: processing initial multi-source data to obtain multi-dimensional time series data; decomposing and reconstructing the multi-dimensional time series data to obtain frequency band multi-dimensional time series data of multiple frequency bands; for each of the frequency band multi-dimensional time series data, establishing a frequency band data regression model and screening a core feature set; based on the core feature set of the to-be-evaluated period, outputting a monitoring data prediction value of the to-be-evaluated period through the frequency band data regression model, and determining abnormal monitoring data based on the monitoring data prediction value and target evaluation data. 2.The time-division fractional band feature engineering-based abnormal data identification method according to claim 1, characterized in that, The processing of the initial multi-source data to obtain the multi-dimensional time series data comprises the following steps: performing integrity check on the obtained initial multi-source data to obtain an integrity check result; the integrity check is realized by checking the meta information of each type of data; performing data outlier cleaning on the initial multi-source data with the integrity check result being complete, and eliminating abnormal data to obtain non-abnormal multi-source data; the data outlier cleaning comprises adopting Mahalanobis distance to detect outliers of multi-dimensional data and adopting Grubbs criterion cycle to detect outliers of single-dimensional data; performing data missing filling on the non-abnormal multi-source data to obtain complete multi-source data; the data missing filling comprises adopting Kriging interpolation method to fill data of spatial data, adopting linear interpolation method to fill data of time data, or adopting inverse distance weighting method to fill data of geographic data; dividing load days based on load data in the complete multi-source data to obtain multiple types of load days and load day monitoring indexes; dividing other data in the complete multi-source data in a daily granularity to obtain multi-source daily granularity data, and aligning the multi-source daily granularity data with the multiple types of load days and the load day monitoring indexes of the corresponding load days to obtain multi-dimensional time series data of a time division day. 3.The time-division fractional band feature engineering-based abnormal data identification method of claim 2, characterized in that, The decomposition and reconstruction of the multi-dimensional time series data to obtain the frequency band multi-dimensional time series data of multiple frequency bands comprises the following steps: performing empirical mode decomposition on each type of load day respectively to obtain connotative modal components and trend items of the multiple types of load days; the connotative modal components comprise multiple frequencies; performing Lempel-Ziv algorithm reconstruction on multiple connotative modal components of each type of load day respectively to obtain high-frequency items and low-frequency items of each type of load day; taking the high-frequency items, the low-frequency items and the trend items of each type of load day as the frequency band multi-dimensional time series data. 4.The time-division fractional band feature engineering-based abnormal data identification method according to claim 3, characterized in that, The empirical mode decomposition comprises the following steps: constructing multi-dimensional time series based on monitoring data and associated data to obtain multi-dimensional time series features; decomposing the multi-dimensional time series features through an empirical mode decomposition algorithm to obtain connotative modal components and trend items of different frequencies; The Lempel-Ziv algorithm reconstruction comprises the following steps: calculating the complexity of the connotative modal components of each type of load day through the Lempel-Ziv algorithm respectively; The components of each type of load day are sequentially added according to the complexity in descending order, and the sum of the first order number of components of each type of load day and the sum of the second order number of components of each type of load day are taken as high-frequency items and low-frequency items respectively; the first order number is the minimum number of components of each type of load day, which is added from the highest complexity of the components of each type of load day, so that the sum of the complexity of the components is greater than the reconstruction threshold; the second order number refers to the number of components other than the first order number. 5.The time-division fractional band feature engineering-based abnormal data identification method according to claim 2, characterized in that, The other data in the complete multi-source data is divided by day granularity to obtain multi-source day granularity data, including: The fine-grained data is down-sampled by the maximum triangle three-bar dimension reduction algorithm to obtain day granularity data; the fine-grained data refers to data with a sampling frequency less than one day; The coarse-grained data is interpolated by the inverse distance weighted difference algorithm to obtain day granularity data; the coarse-grained data refers to data with a sampling frequency greater than one day. 6.The time-division fractional band feature engineering-based abnormal data identification method according to claim 1, characterized in that, The frequency band data regression model is established and the core feature set is screened, including: Feature extraction is performed on the frequency band multi-dimensional time series data to obtain an internal feature sequence and a plurality of external feature sequences; the internal feature sequence is used to represent the correlation change relationship of the monitoring data in the time dimension; the external feature sequence is used to represent the change of other data; The correlation coefficients of the monitoring data sequence with the internal feature sequence and the plurality of external feature sequences are calculated, the plurality of feature sequences are sorted in descending order of the absolute values of the correlation coefficients, and a feature importance sequence is obtained; A frequency band data regression model is constructed, and based on the performance of the frequency band data regression model under a multi-scale importance threshold, the plurality of feature sequences are screened to obtain an initial core feature set; Based on the correlation coefficients between the plurality of feature sequences, the initial core feature set is checked for redundancy, and redundant feature sequences are removed to obtain a core feature set.
7. The method of claim 6, wherein the method further comprises: The plurality of feature sequences are screened to obtain an initial core feature set, including: Based on the current importance threshold, a current core feature set is determined; The current core feature set is divided into a plurality of training sets and a validation set; An initial frequency band data regression model is trained through the plurality of training sets to obtain a trained frequency band data regression model; The trained frequency band data regression model is verified through the validation set to obtain a model performance; A new importance threshold is determined, and the new importance threshold is taken as the current importance threshold, and the performance of the trained frequency band data regression model is repeatedly determined; The model with the best performance is taken as the final frequency band data regression model, the corresponding current core feature set is taken as the initial core feature set, and the current importance threshold is taken as the final importance threshold; the best performance refers to the minimum mean square error of the model. 8.The time-division fractional band feature engineering-based abnormal data identification method of claim 1, wherein, When the abnormal monitoring data is abnormal electricity price data, it further includes processing the abnormal electricity price data to determine the behavior causing the abnormality. 9.The time-division fractional band feature engineering-based abnormal data identification method of claim 8, wherein, The abnormal electricity price data is processed to determine the behavior causing the abnormality, including: determining a category to which the abnormal electricity price data belongs; the category includes a peak day high frequency anomaly, a peak day low frequency anomaly, a peak day trend anomaly, a valley day high frequency anomaly, a valley day low frequency anomaly, a valley day trend anomaly, a flat day high frequency anomaly, a flat day low frequency anomaly, and a flat day trend anomaly; determining one or more category indicators based on the category to which the abnormal electricity price data belongs; matching the category indicators to corresponding indicator ranges to determine the behavior causing the anomaly.
10. A time-of-day and frequency band feature engineering based abnormal data identification system using the time-of-day and frequency band feature engineering based abnormal data identification method according to any one of claims 1-9, characterized in that, The method comprises a data collection and preprocessing module, a data decomposition and reconstruction module, a time-of-day and day-of-week and frequency band feature engineering construction module, and an abnormal data identification module. The data collection and preprocessing module is configured to process initial multi-source data to obtain multi-dimensional time series data. The data decomposition and reconstruction module is configured to decompose and reconstruct the multi-dimensional time series data to obtain frequency band multi-dimensional time series data of multiple frequency bands. The time-of-day and day-of-week and frequency band feature engineering construction module is configured to establish a frequency band data regression model and screen a core feature set for each frequency band multi-dimensional time series data. The abnormal data identification module is configured to output a monitoring data prediction value of an evaluation period to be evaluated by the frequency band data regression model based on the core feature set of the evaluation period to be evaluated, and determine abnormal monitoring data based on the monitoring data prediction value and target evaluation data.
Citation Information
Patent Citations
Power consumption behavior diagnosis system and method based on anomaly detection model
CN119918005A
Frequency converter fault prediction method and system based on machine learning
CN120763808A
Method and apparatus for obtaining power battery life data, computer device, and medium
WO2021208079A1