Method and system for monitoring abnormal behavior data of generator set
By constructing a decomposition-integration prediction model and an autoregressive integral moving average model, the problems of threshold determination and historical anomaly interference in multi-dimensional data anomaly monitoring of generator sets were solved, achieving accurate anomaly identification and monitoring, and improving the safety and stability of the power system.
Patent Information
- Application Number
- CN202511516937.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-23
AI Technical Summary
Existing technologies suffer from difficulties in determining thresholds, interference from historical anomalies, and a lack of self-learning capabilities in multi-dimensional data anomaly monitoring of generator sets, leading to inaccurate monitoring and impacting the safety and stability of the power system.
By constructing a decomposition and integration prediction model based on historical data, and combining structural and behavioral indicators, multi-dimensional data decomposition, reconstruction, and integration prediction are carried out. The maximum mutual information coefficient and autoregressive integral moving average model are used to identify and correct abnormal data and set dynamic monitoring thresholds.
It enables precise monitoring of multi-dimensional data of generator sets, reduces the difficulty and computational complexity of anomaly identification, improves the accuracy and reliability of anomaly identification, and adapts to monitoring needs under complex operating conditions.
Smart Images

Figure CN120995413A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of power generation data monitoring, and specifically discloses a method and system for monitoring abnormal behavior data of a generator set. BACKGROUND
[0002] As the core unit of power production and output of the power system, the multi-dimensional data generated in the behavior process of the generator set is the key basis for reflecting the operation state of the generator set, ensuring the balance and safety and stability of the power supply system. The stable output and reasonable fluctuation of such data are directly related to the load matching accuracy, power transmission efficiency and overall operation safety of the power grid. If the data fluctuates abnormally and is not timely identified and disposed, it may lead to power supply and demand imbalance, deviation of power grid operation parameters from the safety interval, and even affect the stability of power transmission, posing a threat to the reliable operation of the power system. At present, the abnormal monitoring of the above-mentioned multi-dimensional data of the generator set mainly relies on two types of technical paths: one is the monitoring index based on the operation characteristics of the power system, including the related parameters used to evaluate the operation state in the traditional power system; the other is the abnormal identification method combined with intelligent algorithms, such as comprehensive evaluation of the abnormal characteristics of multi-dimensional data through fuzzy decision theory, or construction of an index screening model based on machine learning to realize automatic identification of data abnormalities. In specific research directions, on the one hand, focus is placed on the proposal of new monitoring criteria, such as the coupling relationship between power generation, grid-connected power and power flow through sensitivity analysis of optimal power flow, to assist in identifying data abnormalities; on the other hand, emphasis is placed on optimizing the intelligence level of the identification method to improve the accuracy of abnormal identification. However, the existing abnormal monitoring methods for multi-dimensional data of the generator set still have significant technical shortcomings, making it difficult to meet the precise monitoring needs under complex working conditions.
[0003] Therefore, the present application proposes a method and system for monitoring abnormal behavior data of a generator set, which takes into account the working condition adaptability and cross-dimensional collaborative identification capability, and has the self-optimizing feature of the abnormal monitoring method, to solve the problems of threshold difficulty, historical abnormal interference and lack of self-learning ability in the prior art, ensure the accuracy of the generator set data monitoring and the safety of the power system operation, and determine the threshold value of the pre-evaluation index based on the decomposition and integration prediction method. SUMMARY
[0004] The present application aims to provide a method and system for monitoring abnormal behavior data of a generator set, which solves the problems of threshold difficulty, historical abnormal interference and lack of self-learning ability in the prior art of power data abnormal detection, and the specific scheme is as follows: A method for monitoring abnormal behavior data of a power generating unit, comprising: obtaining historical data of a unit group; the historical data comprising historical monitoring data and historical characteristic data; the historical characteristic data comprising historical meteorological data, historical load data and historical fuel data; based on the historical data, constructing index data, and based on the index data, correcting abnormal values of the historical monitoring data to obtain corrected historical monitoring data; the index data comprising structure indexes and behavior indexes; decomposing and reconstructing the corrected historical data to obtain decomposed historical data of multiple frequencies, based on the decomposed historical data, respectively performing data prediction to obtain decomposed monitoring data prediction values of multiple frequencies, and integrating the decomposed monitoring data prediction values to obtain monitoring data prediction results; the corrected historical data comprising the corrected historical monitoring data and the historical characteristic data; based on the monitoring data prediction results and a preset deviation range of monitoring data, setting a monitoring data range, and based on the monitoring data range, determining whether the monitoring data of the unit is abnormal.
[0005] Further, based on the historical data, constructing structure indexes for the unit group, comprising: determining a share index of the unit group based on a share proportion of the unit group in multiple unit groups; summing up share proportions of a preset number of unit groups with earlier share proportions as a domination index of the unit group; based on the share index and the domination index, determining an abnormal unit group.
[0006] Further, based on the historical data, constructing behavior indexes for each unit in the abnormal unit group, comprising: performing period alignment processing on historical monitoring data of different days to obtain period-aligned monitoring data; screening historical characteristic data to determine similar days of a to-be-tested day; analyzing the period-aligned monitoring data of the similar days using a box plot to obtain a normal data range of the monitoring data; taking monitoring data of the to-be-tested day that is not in the normal data range as abnormal monitoring data; taking a unit with abnormal monitoring data as an abnormal unit, and taking the abnormal monitoring data as abnormal values.
[0007] Further, the screening of the historical feature data to determine the similar day of the to-be-tested day comprises: constructing a to-be-tested day feature vector based on the to-be-tested day feature data; constructing historical operation day feature vectors based on the historical feature data, and combining a plurality of historical operation day feature vectors to obtain a historical operation feature matrix; constructing historical operation day monitoring vectors based on historical monitoring data, and combining a plurality of historical operation day monitoring vectors to obtain a historical monitoring feature matrix; extracting features of a plurality of historical operation days in the historical operation feature matrix to obtain a plurality of historical operation feature vectors corresponding to a plurality of features; calculating the relationship strength between the historical operation feature vector of each feature and the historical monitoring feature matrix by using the maximum mutual information coefficient to obtain a feature weight vector; calculating the feature correlation coefficient of each feature based on the similarity of the to-be-tested day and a plurality of historical operation days on each feature; performing weighted average summation on the feature correlation coefficients based on the feature weight vector to obtain a historical day correlation coefficient of each historical operation day; and screening the historical operation days as the similar days according to the high and low of the historical day correlation coefficients.
[0008] Further, the feature weight vector is: ; wherein, represents the weight of the kth feature; k represents a feature variable, and K represents the total number of features; represents the historical operation feature vector of the kth feature and the historical monitoring feature matrix P. represents the historical operation feature vector; P represents the historical monitoring feature matrix; The calculation formula of the maximum mutual information coefficient is: ; wherein, represents the maximum value under the condition that ; represents that the total number of grids of the historical operation feature vector and the historical monitoring feature matrix under a specific grid division is less than ; represents a function related to the sample size; S represents the sample size; represents the historical operation feature vector of the kth feature and the historical monitoring feature matrix P. represents a logarithmic function with base 2; min represents the minimum value; represents the number of grids after the data is gridded when calculating the mutual information. The feature correlation coefficient is: ; wherein, The eigenvalue represents the eigencorrelation coefficient between the k-th feature of the historical operating day s and the day to be measured; s represents the historical operating day variable. This represents the minimum absolute difference among all characteristics across all historical operating days. This represents the k-th feature value of the day to be measured; This represents the k-th feature value of the s-th historical running day; Represents the resolution coefficient; This represents the maximum absolute difference among all characteristics across all historical running days. The historical daily correlation coefficient is:
[0009] in, This represents the historical day correlation coefficient for the s-th historical operating day.
[0010] Furthermore, the normal data range is as follows: [Q1-1.5(Q3-Q1),Q3+1.5(Q3-Q1)]; Where Q1 is the first quantile parameter; Q3 is the third quantile parameter; The monitoring data for the day to be measured are as follows: ; ; in, This indicates the offset rate of unit i on abnormal historical operating days; This represents the weighted average of the monitoring data for unit i; The value represents the weighted average of monitoring data from units of the same type as unit i during normal operation; h represents the time period variable. This represents the monitoring data of unit i during each time period h; This indicates the amount of monitoring data consumed by unit i during each time period h on abnormal historical operating days.
[0011] Furthermore, the predicted results of the unit monitoring data are obtained, including: performing multivariate empirical mode decomposition on the corrected historical data to obtain multiple intrinsic mode functions and trend terms; reconstructing the multiple intrinsic mode functions into high-frequency terms and low-frequency terms; and using a dynamic model to process the trend terms, the high-frequency terms and the low-frequency terms to obtain the predicted values of the unit monitoring data for each time period.
[0012] Furthermore, constructing the dynamic model includes: using the augmented Dickey-Fuller test to perform unit root tests on the low-frequency and trend terms to determine whether the sequence is stationary, obtaining a stationarity test result; based on the stationarity test result, establishing an autoregressive integral moving average model, including: if the sequence passes the stationarity test, establishing a first autoregressive integral moving average model; if the sequence fails the stationarity test, establishing a second autoregressive integral moving average model after differencing the sequence; the sequence refers to the low-frequency term and / or trend term; based on the autoregressive integral moving average model, according to the self-phase... The parameters of the autoregressive integral moving average model are determined by using the correlation function and partial autocorrelation function to obtain the final autoregressive integral moving average model. The autoregressive conditional heteroscedasticity (ACH) effect of the residuals of the final autoregressive integral moving average model is examined, and the dynamic model is determined based on whether the residuals are correlated. This includes: if the residuals are correlated, a generalized autoregressive conditional heteroscedasticity (GCH) model is constructed for the residual sequence, and the GCH model is used as the dynamic model; if the residuals are not correlated, the residuals of the final autoregressive integral moving average model are used as the dynamic model.
[0013] Furthermore, the autoregressive integral moving average model is as follows: ; in, This represents the value of the monitoring data sequence for each time period on day s. This indicates that the monitoring data sequence for each time period is in s- The value of the day; Indicates the first constant term; Let L represent the number of autoregressive variables; Represents the coefficient of the autoregressive term; represents the residual term on day s; e represents the moving average term variable, and E represents the total number of moving average terms; The coefficient representing the moving average term; This represents the residual term on day se; k represents the coefficient variable, and K represents the total number of coefficient variables; Represents eigenvalues The coefficient; This represents the k-th feature value of the s-th historical running day; The generalized autoregressive conditional heteroscedasticity model is as follows: ; ; in, This represents the conditional standard deviation on day s. Represents the standardized residual; Represents conditional variance. denoted as the second constant term; r represents the autoregressive conditional heteroscedasticity coefficient variable; R represents the total number of autoregressive conditional heteroscedasticity coefficient variables. The coefficients of the autoregressive conditional heteroscedasticity effect variable term; This represents the residual term for day ti introduced from the autoregressive integral moving average model; Let G represent the coefficient variable of the generalized autoregressive conditional heteroscedasticity term, and let G represent the total number of generalized autoregressive conditional heteroscedasticity terms. This represents the coefficient of the generalized autoregressive conditional heteroscedasticity term; This indicates the number of characteristic variables introduced into the variance equation. This represents the total number of features introduced into the variance equation; Indicates the first The coefficients of each feature; Indicates the first The observed values of each feature on day s.
[0014] This invention also provides a system for monitoring abnormal behavior data of generator sets using the aforementioned method, comprising a data collection and preprocessing module, a historical monitoring data anomaly identification and correction module, a decomposition and integration hybrid prediction module, and an early warning module for anomalies; the data collection and preprocessing module is used to acquire historical data of the generator set group; the historical data includes historical monitoring data and historical feature data; the historical feature data includes historical meteorological data, historical load data, and historical fuel data; the historical monitoring data anomaly identification and correction module is used to construct indicator data based on the historical data, and to correct outliers in the historical monitoring data based on the indicator data. The system obtains corrected historical monitoring data; the indicator data includes structural indicators and behavioral indicators; the decomposition and integration hybrid prediction module is used to decompose and reconstruct the corrected historical data to obtain decomposed historical data of multiple frequencies, perform data prediction based on the decomposed historical data to obtain predicted values of decomposed monitoring data of multiple frequencies, and integrate the predicted values of decomposed monitoring data to obtain the monitoring data prediction result; the corrected historical data includes corrected historical monitoring data and historical feature data; the pre-anomaly warning module is used to set the monitoring data range based on the monitoring data prediction result and the preset deviation range of the monitoring data, and determine whether the monitoring data of the unit is abnormal based on the monitoring data range.
[0015] The present invention has the following advantages and beneficial effects: The two-stage abnormality identification correction system designed by the application, in the first stage, the potential abnormal generating units are preliminarily identified through the structure index (the supply-demand matching degree of the associated power generation capacity and on-grid power capacity, the distribution rationality of the clearing price and the clearing capacity), the key monitoring range is quickly narrowed, and the difficulty and the calculation complexity of the overall abnormality identification are effectively reduced; in the second stage, the internal multi-dimensional data of the key units identified in the first stage are further analyzed through the abnormal behavior data index. For example, in combination with the historical data, two kinds of judgment basis of the prior (normal fluctuation range of the power generation capacity / on-grid power capacity based on similar operating conditions) and the posterior (deviation of the clearing price from the normal interval, deviation rate of the power generation capacity from the planned value) are designed, according to the logic of identifying the data abnormal type first and then analyzing the influence of the abnormality on the system, the specific abnormal behavior data item is gradually locked, and finally the multi-dimensional data abnormality of the generating unit is comprehensively and hierarchically accurately identified, so that the abnormal omission or misjudgment caused by the single index monitoring is avoided.
[0016] The historical data abnormality identification correction method designed by the application can identify and correct the hidden abnormality that may exist in the historical monitoring data of the generating unit, and generate clean historical data set without abnormal interference. The correction process effectively alleviates the monitoring deviation caused by directly using the historical data containing abnormality in the traditional method, avoids misjudgment caused by the biased historical data during the current abnormality judgment, and provides a reliable data basis for subsequent model training based on historical data, thereby improving the accuracy of abnormality monitoring and correlation prediction from the source.
[0017] The application deeply correlates the historical characteristic data and the historical monitoring data inside the generating unit, and the decomposed integrated hybrid prediction model constructed by the application can fully adapt to the actual operation scene of the generating unit. On the one hand, the model can accurately predict the normal fluctuation range of the multi-dimensional monitoring data of the generating unit, and provide reasonable and dynamic index threshold (such as the abnormal threshold difference of the power generation capacity under different weather conditions) for abnormality judgment, thereby avoiding the inapplicability of the fixed threshold under complex conditions; on the other hand, the model has higher prediction accuracy for the segmented data (such as the clearing capacity in different time periods or the power generation capacity in different load intervals), can provide more accurate reference benchmark for abnormality monitoring, and further improves the reliability of the overall abnormality identification. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 An exemplary schematic diagram of a system for monitoring abnormal behavior data of a generating unit is provided for the application; Figure 2 An exemplary process for applying the method for monitoring abnormal behavior data of a generating unit to market force abnormality monitoring is provided for the application; Figure 3An exemplary flowchart of applying the method of correcting abnormal values of historical unit monitoring data to market force abnormality monitoring. DETAILED DESCRIPTION
[0019] For the purpose, technical solutions and advantages of the embodiments of the present application to be clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations, and the method of the present application can be used in various abnormality monitoring scenarios, which can not be limited to power generation scenarios only. When the present application is applied to power generation scenarios, the data objects to be monitored can include but are not limited to abnormality detection of online power, power generation, bidding, market force, unit operation time and / or operation state, etc.
[0020] The present application provides a method for monitoring abnormal behavior data of a power generation unit, comprising: Step one, obtaining historical data of a unit group; the historical data includes historical monitoring data and historical feature data; the historical feature data includes historical meteorological data, historical load data and historical fuel data. The historical data refers to multidimensional data related to the historical operation of the unit. The historical data can include historical monitoring data, historical feature data and unit structure data. The monitoring data refers to the data reflecting the internal operation of the unit obtained by monitoring the power generation unit. For example, when monitoring the abnormal behavior data of the power generation unit, the monitoring data can refer to the power generation data of the unit, which can include power generation and online power, etc. For another example, when monitoring the market force abnormal behavior data, the monitoring data can refer to the unit transaction data, which can include unit bidding data, clearing price and clearing quantity, etc. The feature data refers to external feature data affecting the historical monitoring data of the power generation unit. For example, the feature data can include meteorological data, load demand data and fuel supply data, etc. The unit structure data mainly refers to the basic parameters of the unit; the basic parameters of the unit mainly include the type and name of the unit. The unit of each power plant can be regarded as a unit group, and multiple unit groups can be obtained for multiple power plants. The meteorological data mainly refers to data closely related to electricity, such as temperature, humidity, wind speed and irradiance, etc. In some embodiments, the unit structure data can be obtained through the power grid dispatching system, the unit transaction data can be obtained through the power trading platform, the unit power generation data can be obtained through the power generation data acquisition system, the connection with the global reanalysis data set and the regional high-resolution meteorological data can be established, and finally the fuel price related data can be obtained. The fuel price related data includes the electric coal procurement price index and the natural gas procurement comprehensive price, etc.
[0021] In some embodiments, data missing value filling is further included for the acquired historical data, mainly dealing with the missing phenomenon that may exist in the acquisition process of meteorological data and monitoring data. For meteorological data, Kriging interpolation method is used for repair. For the time correlation characteristics of one-dimensional monitoring data of each unit, linear interpolation method is used for filling.
[0022] In step two, based on the historical data, index data is constructed, and based on the index data, the abnormal values of the historical monitoring data are corrected to obtain corrected historical monitoring data; the index data includes structure index and behavior index; the index data includes structure index and abnormal behavior data index. The structure index is used to identify potential abnormal unit groups. For example, potential market force manipulation unit identification is performed. For another example, an abnormal unit group of on-grid power is identified. The behavior index is used to identify potential abnormal units. For example, the report and clearing history data of potential market force manipulation units are identified. For another example, on-grid power abnormality of units is identified.
[0023] In some embodiments, based on the historical data, a structure index of a unit group is constructed, including: Based on the share ratio of the unit group in a plurality of unit groups, a share index of the unit group is determined. When monitoring market force abnormal behavior, the share ratio can be market share. When monitoring unit power generation abnormal behavior data, the share ratio can be the total power generation ratio of the unit. The expression of the share index is: ; Wherein, HHI represents the share index; n represents the unit group variable, N represents the total number of unit groups; represents the monitoring data consumption of unit group n in the total unit group, for example, when monitoring unit power generation abnormal behavior data, it can be the on-grid power of the unit group; Q represents the monitoring data consumption of the total unit group.
[0024] The sum of the share ratios of the top preset number of unit groups in the share ratio is taken as the dominance index of the unit group. For example, the sum of the share ratios of the top 4 unit groups in the on-grid power is taken as the dominance index. Based on the share index and the dominance index, an abnormal unit group is determined. The unit group with both the share index and the dominance index greater than the corresponding threshold value is taken as the abnormal unit group.
[0025] In some embodiments, based on the historical data, a behavior index of each unit in the abnormal unit group is constructed, including: time period alignment processing is performed on the historical monitoring data of different days to obtain time period aligned monitoring data; the historical feature data is screened to determine a similar day of the to-be-tested day; the time period aligned monitoring data of the similar day is analyzed by using a box plot to obtain a normal data range of the monitoring data; the normal data range is: [Q1-1.5(Q3-Q1), Q3+1.5(Q3-Q1)]; wherein Q1 is a first quantile parameter; and Q3 is a third quantile parameter.
[0026] The monitoring data of the to-be-tested day that is not within the normal data range is taken as abnormal monitoring data; the unit that has abnormal monitoring data is taken as an abnormal unit, and the abnormal monitoring data is taken as an abnormal value. The monitoring data of the to-be-tested day is: ; ; wherein, represents the deviation rate of unit i on the abnormal historical operation day; represents the weighted average value of the monitoring data of unit i, for example, the weighted average value of the on-grid power; represents the weighted average value of the monitoring data of the unit of the same type when the unit is normally operated; h represents a time period variable, represents the monitoring data of unit i in each time period h; represents the monitoring data consumption of unit i in each time period h on the abnormal historical operation day.
[0027] In some embodiments, the historical feature data is screened to determine similar days of the to-be-tested day, comprising: constructing a to-be-tested day feature vector based on the to-be-tested day feature data; constructing a historical operation day feature vector based on the historical feature data, and combining a plurality of historical operation day feature vectors to obtain a historical operation feature matrix; constructing a historical operation day monitoring vector based on the historical monitoring data, and combining a plurality of historical operation day monitoring vectors to obtain a historical monitoring feature matrix; extracting the features of a plurality of historical operation days in the historical operation feature matrix to obtain a plurality of historical operation feature vectors corresponding to a plurality of features; calculating the relationship strength between the historical operation feature vector of each feature and the historical monitoring feature matrix by using the maximum mutual information coefficient to obtain a feature weight vector; the feature weight vector is: ; wherein, represents the weight of the kth feature; k represents a feature variable, and K represents the total number of features; represents the historical operation feature vector of the kth feature and the maximum mutual information coefficient of the historical monitoring feature matrix P; represents the historical operation feature vector; P represents the historical monitoring feature matrix.
[0028] The calculation formula of the maximum mutual information coefficient is: ; wherein, represents the maximum value under the condition that ; represents that the total number of grids of the historical operation feature vector and the historical monitoring feature matrix under a certain grid division is less than ; represents a function about sample size; S represents sample size; represents the historical operation feature vector of the kth feature and the mutual information of the historical monitoring feature matrix P, which measures how much the uncertainty of one variable is reduced after knowing the value of another variable; represents the logarithmic function with base 2; min represents taking the minimum value; represents taking the absolute value.
[0029] Based on the similarity of each feature between the to-be-tested day and the plurality of historical operation days, a feature correlation coefficient of each feature is respectively calculated; the feature correlation coefficient is: ; wherein, represents the feature correlation coefficient of the kth feature between the to-be-tested day and the historical day s, which measures the shape similarity of the to-be-tested day and the historical day s on a single feature k; the closer the value is to 1, the more similar the shapes are; s represents the historical operation day variable; represents the minimum absolute difference in all historical operation days and all features; represents the kth feature value of the to-be-tested day; represents the kth feature value of the s th historical operation day; represents a resolution coefficient, which is usually 0.5, and is used to adjust the resolution of the calculation result; the smaller the value is, the greater the resolution is; represents the maximum absolute difference in all historical operation days and all features.
[0030] The feature correlation coefficients are weighted and averaged to obtain a historical day correlation coefficient of each historical operation day; the historical day correlation coefficient is:
[0031] wherein, represents the historical day correlation coefficient of the s th historical operation day.
[0032] The historical operation days are screened according to the high and low of the historical day correlation coefficients as similar days.
[0033] Embodiment 1 Taking monitoring of market force abnormal behavior data as an example, as shown in Figure 2 and Figure 3 , step two is further illustrated: Structural indicators can include share index and share percentage. The share index can be the sum of the squares of the percentages of all cleared bids from each power generation company's units relative to the total cleared market volume. The parameter Q in the share index can be the total cleared volume in the electricity market. This can be viewed as the clearing volume of all units within the nth enterprise; the larger the HHI value, the higher the market concentration and the higher the degree of monopoly. If the value is 1 / N, it indicates that the market share of all enterprises in the industry is the same; if the value is 1, it indicates that the industry is monopolized by a single enterprise. The expression for market share can be: ; Here, V represents the market share of the top M power generation companies in the market, i.e., the clearing volume ratio, which is the core indicator for measuring market concentration. m represents the unit group variable M, which is a set threshold for the number of suppliers. It is usually taken as 4, that is, the total market share of the four largest power suppliers in the market. This indicator judges the market structure by setting a specific threshold (such as Top-4 ≥ 65%): exceeding the threshold indicates that the market has oligopolistic characteristics, and the higher the value, the stronger the concentration. Based on the above structural primary indicator values, potential market-controlling units are identified, and key observation targets are determined.
[0034] Behavioral indicators can be categorized into historical bid normal range and clearing price deviation rate based on pre- and post-event considerations and the transmission relationship between indicators. The historical bid normal range is a pre-event indicator, while the clearing price deviation rate is a post-event indicator. The former is used to identify abnormal generating units bidding high prices, and then the clearing prices of these abnormal units are further anomaly identified. Since the number of segmented bids for the same generating unit varies across different days, granular alignment of the unit's segmented bid data is necessary before calculating the secondary indicators for bid anomalies. The specific steps for granular alignment are as follows: Assume that unit 1 submits its segmented bid on day s. Segmentation start segment and segment termination segment They are respectively:
[0035]
[0036]
[0037] Assume that unit 1 is in the The segmented pricing for the day is The starting segment is The segment termination segment is First, the application segments for generating units were standardized, and the granularity of the standardized pricing segments was then determined as follows: Tianhe The union of the daily price segments, i.e. The union processing sorts all the segmentation points, i.e. the new offer segmentation is . The offer of the unit 1 on the s-th day and the day is re-disassembled according to the new offer segmentation, and the segmentation does not have a fill-in maximum value, such as 99999, which is ignored in the index calculation.
[0038] The normal range of the historical offer is determined based on the similar day, and then the normal range and abnormal value of the offer are determined by analyzing the offer data of the similar day by using a box plot.
[0039] The expression for determining the similar day is as follows: MIC-GRA similar day screening method: assuming that there are S historical operation days, each historical day is composed of K features and a target offer sequence, a feature matrix with a shape of is constructed, each row represents a historical day, and each column represents a feature, and the feature matrix is denoted as: ; wherein, is the k-th feature value of the s-th historical day. The feature vector of the to-be-tested day also has a shape of .
[0040] It is assumed that the target offer sequence is P, which has a shape of , wherein H is the number of segments of the unit segmentation offer, and each row represents a complete offer curve corresponding to a historical day: ; wherein, is the value of the unit in the h-th segment of the s-th historical day (historical operation day monitoring vector); and P is the historical monitoring feature matrix.
[0041] Constructing a feature weight: using the maximal information coefficient (MIC) to calculate the strength of the relationship between each feature and the offer curve as the weight of the feature: is the vector composed of the k-th feature of all historical days, i.e. the k-th column of the feature matrix. P is the set of offer sequences of all historical days.
[0042] After calculating the MIC value of each feature column, the feature weight vector is obtained by normalization. The similarity between the to-be-tested day and each historical day in all features is calculated, and the correlation coefficient of each feature is calculated after standardizing the features, i.e. dividing each feature value by the mean value of the historical day and the to-be-tested day. The historical day correlation coefficient is calculated. The obtained historical day correlation coefficient is sorted from large to small, and the larger the coefficient is, the greater the similarity between the two is. The dates with high rankings are selected as the similar days.
[0043] The method for determining abnormal values by using a box plot is mainly based on a quartile range method. The principle of the quartile range is as follows: abnormal range determination: the segmented bidding sequences of the historical similar day and the day to be tested are sorted according to each segment, and after sorting, the bid values of each segment are divided into four equal parts, denoted as Q1, Q2, Q3, which represent the first, second and third quartiles, respectively. The expression of the quartile range is IQR = Q3-Q1, and Q1-1.5IQR to Q3+1.5IQR is set as the non-abnormal range, and the data outside the range is the abnormal behavior data.
[0044] The unit that bids too high is identified through the historical bidding normal range, and then the post-index calculation is performed. In the uniform processing of the bidding granularity, the bid filled with 99999 due to no bidding in the segment is not recorded as an abnormal bid.
[0045] The monitoring data of the day to be tested can be the clearing price deviation rate, and the clearing price deviation rate is taken as the post-index to determine the influence of the high bidding behavior on the clearing price. Among them, is the clearing price deviation rate of the high bidding unit i on the high bidding day (abnormal day) is the weighted average value of the clearing price of the high bidding unit; is the weighted average value of the clearing price of the high bidding unit; is the weighted average value of the clearing price of the normal day of the same type unit. is the clearing price of the abnormal bidding unit (such as the unit that bids too high) in each period h, is the winning amount of the high bidding unit in each period h on the high bidding day. The deviation rate of the clearing price greater than the threshold value indicates that the unit not only bids high, but also the high bidding behavior has an impact on the clearing price.
[0046] After determining the abnormal bidding data based on the above index, the abnormal values are corrected. The segmented bidding abnormal values are corrected to the average value of the normal bidding of the same type unit on the same day. If there is no normal bidding of the same type unit on the same day, the average value of the normal bidding of the same type unit on the historical similar day is corrected. Thus, the corrected historical unit transaction data is obtained.
[0047] Step three, decompose and reconstruct the corrected historical data to obtain decomposed historical data of multiple frequencies, respectively perform data prediction based on the decomposed historical data to obtain decomposed monitoring data prediction values of multiple frequencies, and integrate the decomposed monitoring data prediction values to obtain a monitoring data prediction result; the corrected historical data includes corrected historical monitoring data and historical feature data.
[0048] In some embodiments, obtaining the unit group monitoring data prediction result comprises: performing multivariate empirical mode decomposition on the corrected historical data to obtain a plurality of intrinsic mode functions and a trend item; reconstructing the plurality of intrinsic mode functions into a high-frequency item and a low-frequency item; and processing the trend item, the high-frequency item and the low-frequency item by using a dynamic model to obtain a monitoring data prediction value of the unit group at each time period.
[0049] In some embodiments, the plurality of intrinsic mode functions (IMFs) after decomposition are reconstructed into a high-frequency item and a low-frequency item based on a Lempel-Ziv algorithm, and the specific steps of the Lempel-Ziv algorithm are as follows: The lz value of each IMF is calculated, and when the cumulative value of the IMFs exceeds a threshold value, the corresponding IMF is the high-frequency item, and the threshold value is generally 0.8, and specifically:
[0050] wherein z represents the zth IMF after sorting according to the frequency from high to low; represents the number of IMFs after sorting according to the frequency from high to low and satisfying the above formula; represents the complexity of the zth IMF; represents a preset threshold value, and is generally 0.8. If the above formula is satisfied, to the sum of the high-frequency item and the low-frequency item is the high-frequency item, and the sum of the remaining IMFs is the low-frequency item, and the trend item represents the trend of the sequence.
[0051] In some embodiments, the dynamic model is constructed by: unit root test on the low-frequency item and the trend item by using an augmented Dickey-Fuller test method to determine whether the sequence is stationary to obtain a stationary test result. For example, unit root test on the low-frequency item and the trend item by using an augmented Dickey-Fuller test (ADF) method to determine whether the sequence is stationary, if the sequence does not contain a unit root as tested by the ADF method, i.e., passes the stationarity test, an ARIMA(L, 0, E) model can be directly established; if the sequence contains a unit root as tested by the ADF method, i.e., fails the stationarity test, the sequence needs to be differenced, and an ARIMA(L, d, E) model is established.
[0052] Based on the stationary test result, an autoregressive integrated moving average model is established, comprising: if the sequence passes the stationarity test, a first autoregressive integrated moving average model is established; if the sequence fails the stationarity test, a second autoregressive integrated moving average model is established after the sequence is differenced; the sequence refers to the low-frequency item and / or the trend item; and the autoregressive integrated moving average model is: wherein, denotes the value of the monitoring data sequence of each period on day s, denotes the value of the monitoring data sequence of each period on day s, denotes the value of the monitoring data sequence of each period on day s; denotes the first constant term; denotes the autoregressive term variable, and L denotes the total number of autoregressive terms; denotes the coefficient of the autoregressive term; denotes the residual term on day s, e denotes the moving average term variable, and E denotes the total number of moving average terms; denotes the coefficient of the moving average term; denotes the residual on day s-e, k denotes the coefficient variable, and K denotes the total number of coefficient variables; denotes the coefficient of the characteristic value ; denotes the kth characteristic value of the s th historical operation day.
[0053] Based on the autoregressive integrated moving average model, parameters of the autoregressive integrated moving average model are determined according to the autocorrelation function and the partial autocorrelation function, and a final autoregressive integrated moving average model is obtained; The autoregressive conditional heteroscedasticity effect of the final autoregressive integrated moving average model residual is tested, and based on whether the residual has correlation, the dynamic model is determined, including: if the residual has correlation, a generalized autoregressive conditional heteroscedasticity model is constructed for the residual sequence, and the generalized autoregressive conditional heteroscedasticity model is taken as the dynamic model; if the residual has no correlation, the final autoregressive integrated moving average model residual is taken as the dynamic model. The generalized autoregressive conditional heteroscedasticity model is: ; ; wherein, denotes the conditional standard deviation of day s, denotes the standardized residual; denotes the conditional variance, denotes the second constant term; r denotes the autoregressive conditional heteroscedasticity effect coefficient variable, and R denotes the total number of autoregressive conditional heteroscedasticity effect coefficient variables; is the coefficient of the autoregressive conditional heteroscedasticity effect coefficient variable term; denotes the residual term of day t-i introduced from the autoregressive integrated moving average model; denotes the coefficient variable of the generalized autoregressive conditional heteroscedasticity term, and G denotes the total number of generalized autoregressive conditional heteroscedasticity terms; denotes the coefficient of the generalized autoregressive conditional heteroscedasticity term; denotes the characteristic number variable introduced in the variance equation, and these introduced variance characteristics are a subset of all characteristics, denotes the total number of features introduced in the variance equation; denotes the coefficient of the th feature; denotes the observation of the th feature on the s-th day.
[0054] Embodiment 2 Step three is further illustrated by taking monitoring market power abnormal behavior data as an example: Based on the decomposition-integration prediction framework, the modified historical unit transaction data is decomposed and reconstructed into different frequency modules, the similar days are determined based on the features to obtain the high-frequency module transaction data prediction values, the medium and low-frequency modules are modeled to obtain the prediction values of the medium and low-frequency modules, and finally the prediction values of each decomposition module are integrated to obtain the final quotation data prediction result.
[0055] In order to reduce the calculation cost, this step only predicts the quotation of the units identified by the two-stage index, i.e. the units with market power abuse ability are predicted, and the threshold of quotation abnormality is determined. The units described in step three are all units with market power abuse ability.
[0056] The decomposition-integration prediction framework mainly includes: a) The multi-dimensional data decomposition method multi-component empirical mode decomposition (MEMD) is used to decompose the unit quotation data and unit quantity data into intrinsic mode functions (IMF) and trend items of different frequencies. The multi-component empirical mode decomposition process is as follows: after the two-stage index is modified, the historical data of the segmented quotation of the unit (modified historical unit monitoring data) is obtained, each segment of the quotation sequence is combined with the feature sequence to form a multi-sequence , wherein is the quotation history sequence of the i-th unit h-th segment, the multi-sequence is input into the MEMD method, and a plurality of IMFs are output: and trend item The decomposed multi-sequence has the same row and column number as the original sequence.
[0057] b) The high-frequency, low-frequency and trend items after multi-component decomposition are in the form of: , wherein represent the high-frequency item, the low-frequency item and the trend item of the quotation sequence, High, low and trend components of the characteristic sequence respectively. The historical similar days of the day to be predicted are determined by using the meteorological features, fuel price features and load demand features in the high frequency component. Based on the MIC-GRA similar day screening method described above, the mean value of the high frequency data of the similar days after the similar day set is determined, and the similar day name unit is corrected to obtain the prediction value of the high frequency component of the day to be predicted.
[0058] c) Constructing the quotation ARIMA-GARCH model for the low frequency module and the trend component respectively: Determine the ARIMA model parameters according to the ACF (Autocorrelation Function) and the PCAF (Partial Autocorrelation Function) and . Among them, the represents the value of each segmented quotation sequence in s days.
[0059] Test the ARCH effect of the quotation ARIMA model residual : analyze the correlation of the residual, if there is no correlation, the modeling is ended and step 5 is entered; if there is correlation, the quotation GARCH model is established. The prediction is obtained by using the model, and the prediction value of the low frequency component and the segmented quotation of the low frequency component is obtained.
[0060] d) Integrate the prediction values of the high frequency, low frequency and trend components of each segmented quotation to obtain the final quotation prediction value of each unit in each segment.
[0061] Step four, based on the unit monitoring data prediction result and the preset deviation range of the unit monitoring data, set the unit monitoring data range, and determine whether the current unit monitoring data is abnormal based on the unit monitoring data range. For example, the prediction value of the power monitoring data is set as the standard value of the unit segmented monitoring data, the reasonable monitoring data range of each unit is determined based on the human preset deviation range, and the pre-unit abnormal identification warning is realized.
[0062] As Figure 1As shown, the application also provides a system for monitoring abnormal behavior data of a generator set, comprising a data collection preprocessing module, a historical monitoring data abnormality identification correction module, a decomposition integrated hybrid prediction module and an abnormality identification module; the data collection preprocessing module is used for obtaining historical data of a group of units; the historical data comprises historical monitoring data and historical characteristic data; the historical characteristic data comprises historical meteorological data, historical load data and historical fuel data; the historical monitoring data abnormality identification correction module is used for constructing index data based on the historical data, and correcting abnormal values of the historical monitoring data based on the index data to obtain corrected historical monitoring data; the index data comprises structure indexes and behavior indexes; the decomposition integrated hybrid prediction module is used for decomposing and reconstructing the corrected historical data to obtain decomposition historical data of multiple frequencies, respectively performing data prediction based on the decomposition historical data to obtain decomposition monitoring data prediction values of multiple frequencies, and integrating the decomposition monitoring data prediction values to obtain monitoring data prediction results; the corrected historical data comprises the corrected historical monitoring data and the historical characteristic data; the abnormality identification module is used for setting a monitoring data range based on the monitoring data prediction results and a preset deviation range of monitoring data, and determining whether monitoring data of a unit is abnormal based on the monitoring data range.
[0063] The above merely describes the preferred embodiments of the present application and is not used to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for monitoring abnormal behavior data of generator sets, characterized in that, include: Acquire historical data of the unit group; the historical data includes historical monitoring data and historical characteristic data; the historical characteristic data includes historical meteorological data, historical load data, and historical fuel data; Based on the historical data, indicator data is constructed, and based on the indicator data, outliers in the historical monitoring data are corrected to obtain corrected historical monitoring data; the indicator data includes structural indicators and behavioral indicators. The corrected historical data is decomposed and reconstructed to obtain decomposed historical data of multiple frequencies. Data prediction is performed based on the decomposed historical data to obtain decomposed monitoring data prediction values of multiple frequencies. The decomposed monitoring data prediction values are then integrated to obtain the monitoring data prediction results. The corrected historical data includes corrected historical monitoring data and historical feature data; Based on the predicted results of the monitoring data and the preset deviation range of the monitoring data, a monitoring data range is set, and based on the monitoring data range, it is determined whether the monitoring data of the unit is abnormal.
2. The method for monitoring abnormal behavior data of generator sets according to claim 1, characterized in that, Based on the aforementioned historical data, structural indicators for the unit group construction include: The share index of a unit group is determined based on its share among multiple unit groups. The sum of the share percentages of the top-ranked pre-defined number of unit groups is used as the dominance index of the unit group. Based on the share index and the dominance index, the abnormal unit group is determined.
3. The method for monitoring abnormal behavior data of generator sets according to claim 2, characterized in that, Based on the historical data, behavioral indicators are constructed for each unit in the abnormal unit group, including: Historical monitoring data from different days are aligned by time period to obtain time-aligned monitoring data. Historical characteristic data is filtered to identify similar days to the day to be measured; Box plots were used to analyze the monitoring data after alignment of similar day periods to obtain the normal data range of the monitoring data; Monitoring data for the test day that is outside the normal data range will be considered abnormal monitoring data. Units with abnormal monitoring data are designated as abnormal units, and the abnormal monitoring data is designated as abnormal values.
4. The method for monitoring abnormal behavior data of generator sets according to claim 3, characterized in that, The process of filtering historical feature data to determine similar days to the day to be measured includes: Based on the feature data of the day to be measured, a feature vector of the day to be measured is constructed; Based on historical feature data, a feature vector of historical operating days is constructed, and multiple feature vectors of historical operating days are combined to obtain a historical operating feature matrix; Based on historical monitoring data, a monitoring vector for historical operating days is constructed, and multiple historical operating day monitoring vectors are combined to obtain a historical monitoring feature matrix; Features of multiple historical operating days are extracted from the historical operating feature matrix to obtain multiple historical operating feature vectors corresponding to various features; The relationship strength between the historical operating feature vector and the historical monitoring feature matrix for each feature is calculated using the maximum mutual information coefficient, thus obtaining the feature weight vector; Based on the similarity between the day to be tested and multiple historical operating days for each feature, the feature correlation coefficient for each feature is calculated. The historical daily correlation coefficient for each historical operating day is obtained by weighted average summation of the feature correlation coefficients based on the feature weight vector. Historical running days are selected as similar days based on their correlation coefficients.
5. The method for monitoring abnormal behavior data of generator sets according to claim 4, characterized in that, The feature weight vector is: ; in, This represents the weight of the k-th feature; k represents the feature variable, and K represents the total number of features; This represents the historical feature vector of the k-th feature. The maximum mutual information coefficient between the feature matrix P and the historical monitoring feature matrix; P represents the historical operational feature vector; P represents the historical monitoring feature matrix. The formula for calculating the maximum mutual information coefficient is as follows: ; in, Indicates that when the condition is met The maximum value under the given conditions; This indicates that the total number of cells in the historical operation feature vector and historical monitoring feature matrix under a specific grid division must be less than [a certain value]. ; The expression represents a function of the sample size; S represents the sample size. This represents the historical feature vector of the k-th feature. Mutual information with the historical monitoring feature matrix P; This represents a logarithmic function with a base of 2; min indicates taking the minimum value. This refers to the number of grids after data gridding when calculating mutual information; The feature correlation coefficient is: ; in, The eigenvalue represents the eigencorrelation coefficient between the k-th feature of the historical operating day s and the day to be measured; s represents the historical operating day variable. This represents the minimum absolute difference among all characteristics across all historical operating days. This represents the k-th feature value of the day to be measured; This represents the k-th feature value of the s-th historical running day; Represents the resolution coefficient; This represents the maximum absolute difference among all characteristics across all historical running days. The historical daily correlation coefficient is: in, This represents the historical day correlation coefficient for the s-th historical operating day.
6. The method for monitoring abnormal behavior data of generator sets according to claim 3, characterized in that, The normal data range is: [Q1-1.5(Q3-Q1),Q3+1.5(Q3-Q1)]; Where Q1 is the first quantile parameter; Q3 is the third quantile parameter; The monitoring data for the day to be tested are as follows: ; ; in, This indicates the offset rate of unit i on abnormal historical operating days; This represents the weighted average of the monitoring data for unit i; The value represents the weighted average of monitoring data from units of the same type as unit i during normal operation; h represents the time period variable. This represents the monitoring data of unit i during each time period h; This indicates the amount of monitoring data consumed by unit i during each time period h on abnormal historical operating days.
7. The method for monitoring abnormal behavior data of generator sets according to claim 1, characterized in that, The predicted results obtained from the unit monitoring data include: Multivariate empirical mode decomposition is performed on the corrected historical data to obtain multiple intrinsic mode functions and trend terms; Multiple intrinsic mode functions are reconstructed into high-frequency and low-frequency terms; The trend term, the high-frequency term, and the low-frequency term are processed using a dynamic model to obtain the predicted values of the monitoring data of the unit in each time period.
8. The method for monitoring abnormal behavior data of generator sets according to claim 7, characterized in that, Constructing the dynamic model includes: The augmented Dickey-Fuller test method was used to perform unit root tests on low-frequency and trend terms to determine whether the series was stationary, and the stationarity test results were obtained. Based on the stationarity test results, an autoregressive integral moving average model is established, including: if the sequence passes the stationarity test, a first autoregressive integral moving average model is established; if the sequence fails the stationarity test, a second autoregressive integral moving average model is established after differencing the sequence; the sequence refers to the low-frequency term and / or the trend term. Based on the autoregressive integral moving average model, the parameters of the autoregressive integral moving average model are determined according to the autocorrelation function and the partial autocorrelation function, thus obtaining the final autoregressive integral moving average model. The process involves examining the autoregressive conditional heteroscedasticity (ACH) effect of the residuals in the final autoregressive integral moving average model, and determining the dynamic model based on whether the residuals are correlated. This includes: if the residuals are correlated, constructing a generalized autoregressive conditional heteroscedasticity (GCH) model for the residual sequence and using the GCH model as the dynamic model; if the residuals are not correlated, using the residuals of the final autoregressive integral moving average model as the dynamic model.
9. The method for monitoring abnormal behavior data of generator sets according to claim 8, characterized in that, The autoregressive integral moving average model is as follows: ; in, This represents the value of the monitoring data sequence for each time period on day s. This indicates that the monitoring data sequence for each time period is in s- The value of the day; Indicates the first constant term; Let L represent the number of autoregressive variables; Represents the coefficient of the autoregressive term; represents the residual term on day s; e represents the moving average term variable, and E represents the total number of moving average terms; The coefficient representing the moving average term; This represents the residual term on day se; k represents the coefficient variable, and K represents the total number of coefficient variables; Represents eigenvalues The coefficient; This represents the k-th feature value of the s-th historical running day; The generalized autoregressive conditional heteroscedasticity model is as follows: ; ; in, This represents the conditional standard deviation on day s. Represents the standardized residual; Represents conditional variance. denoted as the second constant term; r represents the autoregressive conditional heteroscedasticity coefficient variable; R represents the total number of autoregressive conditional heteroscedasticity coefficient variables. The coefficients of the autoregressive conditional heteroscedasticity effect variable term; This represents the residual term for day ti introduced from the autoregressive integral moving average model; Let G represent the coefficient variable of the generalized autoregressive conditional heteroscedasticity term, and let G represent the total number of generalized autoregressive conditional heteroscedasticity terms. This represents the coefficient of the generalized autoregressive conditional heteroscedasticity term; This indicates the number of characteristic variables introduced into the variance equation. This represents the total number of features introduced into the variance equation; Indicates the first The coefficients of each feature; Indicates the first The observed values of each feature on day s.
10. A system for monitoring abnormal behavior data of a generator set using the method for monitoring abnormal behavior data of a generator set according to any one of claims 1-9, characterized in that, It includes a data collection and preprocessing module, a historical monitoring data anomaly identification and correction module, a decomposition and integration hybrid prediction module, and an anomaly identification module; The data collection and preprocessing module is used to acquire historical data of the unit group; the historical data includes historical monitoring data and historical characteristic data; the historical characteristic data includes historical meteorological data, historical load data, and historical fuel data; The historical monitoring data anomaly identification and correction module is used to construct indicator data based on the historical data, and to correct the outliers in the historical monitoring data based on the indicator data to obtain corrected historical monitoring data; the indicator data includes structural indicators and behavioral indicators. The decomposition-integration-hybrid prediction module is used to decompose and reconstruct the corrected historical data to obtain decomposed historical data of multiple frequencies. Based on the decomposed historical data, data prediction is performed to obtain decomposed monitoring data prediction values of multiple frequencies. The decomposed monitoring data prediction values are then integrated to obtain the monitoring data prediction result. The corrected historical data includes corrected historical monitoring data and historical feature data; The anomaly identification module is used to set the monitoring data range based on the monitoring data prediction results and the preset deviation range of the monitoring data, and to determine whether the monitoring data of the unit is abnormal based on the monitoring data range.
Citation Information
Patent Citations
Medium and long term runoff prediction method based on LSFL combination model
CN111523644A
Tunnel health monitoring abnormal data dynamic early warning method based on ARIMA model
CN116163807A
Economic index abnormity alarm system based on intelligent data analysis
CN119151331A
Short-term photovoltaic power generation prediction method and system based on multi-period clustering and signal reconstruction
CN120450116A
Method for handling multidimensional data
WO2018111116A2