Method and system for monitoring abnormal electricity consumption of user side based on data analysis
By analyzing the periodic dependence of the user-side power consumption data and building a compensation set, the impact of data loss on monitoring results is solved, high-quality compensation of data and accuracy of abnormal detection is achieved, and the system's adaptability and monitoring effect are improved.
Patent Information
- Application Number
- CN202510247159.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-04
AI Technical Summary
In the existing abnormal power monitoring systems on the user side, data loss problems are common. The traditional missing value filling method ignores the time dependence and data correlation, resulting in deviations in filling results and affecting the effectiveness of abnormal detection.
By analyzing the periodic dependence of the real-time power consumption data of the user side, setting the upper limit of the window Tmax, filtering out a complete window without missing values, building a compensation set, dynamically selecting appropriate missing compensation means, generating real-time complete data points, and using density analysis to judge abnormal points, dynamically adjusting the nearest neighbor number K.
Effectively eliminate the impact of data loss on monitoring results, ensure data integrity and accuracy, improve the accuracy of abnormal electricity detection and system adaptability, and reduce the possibility of misjudgment and misjudgment.
Smart Images

Figure CN120179953A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of anomaly monitoring, and particularly to a method and system for monitoring abnormal power consumption of a user terminal based on data analysis. Background Art
[0002] With the rapid development of information technology and intelligent applications, smart grids and energy management systems have become an important part of energy efficiency optimization in modern cities and households. Especially driven by smart meters and Internet of Things technologies, monitoring abnormal power consumption at the user terminal has become an important research direction in the field of energy management. This field focuses on how to timely identify abnormal power consumption situations at the user terminal through real-time monitoring and data analysis, such as sudden increase in power load, equipment failure, or power waste. However, when collecting power consumption data at the user terminal, problems such as data missing and noise interference are often faced. If these problems are not reasonably processed, it will directly affect the analysis effect of subsequent algorithms. If missing values are randomly filled or filled without discrimination in the data preprocessing stage, it may have an adverse impact on the effectiveness and accuracy of the algorithms, and further lead to the failure of anomaly detection.
[0003] Currently, in the user terminal abnormal power consumption monitoring system, the data missing problem is a common challenge. Usually, the generation of missing values may be caused by various factors such as sensor failure, network problems, or data transmission interruption. Although common missing value filling methods, such as mean filling, linear interpolation, forward filling, etc., are widely used, these filling methods are not always suitable for all situations, especially when dealing with complex time series data. Specifically, these methods often ignore the time dependence and the correlation between data, and are prone to cause deviation in the filling results. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a method and system for monitoring abnormal power consumption of a user terminal based on data analysis, and solves the problems in the above background art.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for monitoring abnormal power consumption of a user terminal based on data analysis, comprising the following steps:
[0006] S1: Monitor the real-time power consumption data of the user terminal through a sensor group, analyze the periodic dependence degree of the sequence in combination with the time series property, so as to set the window upper limit T at each data point max , and based on the historical data group, screen the windows without missing data from the window upper limit T at each data point max , and mark them as complete windows;
[0007] S2: Set different missing ranges according to the complete window to perform the missing compensation means test operation, and select corresponding missing compensation means according to different missing conditions to construct a compensation set;
[0008] S3: After preprocessing the real-time power consumption data, combine the window upper limit T at the data points obtained in S1 and S2 max and the compensation set, and take missing compensation means for the real-time data points to generate real-time complete data points;
[0009] S4: Preset the number of neighbors K in advance, determine the K neighbor points adjacent to the real-time complete data points, analyze and compare the density between the real-time complete data points and their neighbor points, judge abnormal points, monitor the abnormal power consumption of the user side, and dynamically adjust the number of neighbors K.
[0010] Preferably, the specific steps of S1 include:
[0011] S11: Use the sensor group to monitor the real-time power consumption data of the user side, generate a power consumption feature vector based on the real-time power consumption data; and store the real-time user data in the cloud database at regular intervals according to the time series to generate a historical data group;
[0012] S12: Based on the historical data group, analyze the periodic dependence between time series data to obtain the lag correlation coefficient Z(k), which is obtained specifically in the following way:
[0013]
[0014] In the formula, X t is the power consumption feature vector at the data point t; X avg is the mean value of the power consumption feature vector; X t+k is the power consumption feature vector at the data point t + k; T is the length of the historical data points in the historical data group; t is the number of the data point; k is the lag step.
[0015] Preferably, the specific steps of S1 also include:
[0016] S13: Based on the lag correlation coefficient Z(k) obtained in S12, draw a relationship graph, and according to the relationship graph, obtain the change trend of the time series at different lag steps k, and based on the change trend of the time series at different lag steps k, set the window upper limit T at each data point max , and the specific setting steps are as follows:
[0017] S131: Based on the relationship graph, determine the decline rate Jz between each data point and the data points at different lag steps k, and compare the adjacent two groups of decline rates Jz in the relationship graph to obtain the decline difference Jc;
[0018] S132. Statistically determine the maximum drop difference Jc max , and use the position corresponding to the maximum drop difference Jc max in the relationship diagram as the upper limit of the lag point. Based on the upper limit of the lag point, determine the window upper limit T at the corresponding data point max .
[0019] Preferably, the specific steps of S1 further include:
[0020] S14. Extract the window upper limit T at several data points from the historical data set max , and screen the windows without missing data from the window upper limit T at each data point max and mark them as complete windows.
[0021] Preferably, the specific steps of S2 include:
[0022] S21. Set different missing ranges for the complete windows. The missing range refers to the range where there are missing values in the windows corresponding to each data point, and perform missing compensation means testing operations on the missing values in the corresponding data points according to different missing conditions;
[0023] S22. During the process of performing the missing compensation means testing operation, determine the filling results of the missing values by various missing compensation means under different missing conditions to obtain the filling value BT. Based on the filling value BT, identify the corresponding missing compensation means under different missing conditions. The specific identification steps are:
[0024] S221. Set different missing conditions, specifically: Among them, Tcb t is the missing value ratio at the data point t; Qc t is the missing value range within the window upper limit T at the data point t max ; T max,t is the window upper limit T at the data point t max ;
[0025] S222. Obtain the deviation value PT, and the deviation value PT is obtained through the following formula: PT = |BT - SZ|, where SZ is the actual value;
[0026] S223. Statistically select the minimum deviation value PT min , and use the missing compensation means corresponding to the minimum deviation value PT min as the best missing compensation means for the missing conditions set in the complete window at the corresponding data point to construct a compensation set. The compensation set includes the best missing compensation means for different missing conditions set in the complete window at the corresponding data point.
[0027] Preferably, the specific steps of S3 include:
[0028] S31. Preprocess the real-time electricity consumption data obtained in S1 to identify the missing values in the real-time electricity consumption data, determine the range of the missing values, and combine the window upper limit T at the data points obtained in S1 and S2 max and the compensation set, and adopt corresponding optimal missing compensation means for the data points corresponding to the real-time electricity consumption data to generate real-time complete data points.
[0029] Preferably, the specific steps of S4 include:
[0030] S41. Preset the number of neighbors K, respectively determine the Euclidean distances between each historical data point in the historical data group and the real-time complete data point, and extract K groups of historical data points adjacent to the real-time complete data point from them to obtain a local data group;
[0031] S42. Based on the local data group, analyze and compare the density between the real-time complete data point and its neighbor points to obtain the anomaly index YZ(p) at the real-time complete data point, which is specifically obtained through the following formula:
[0032]
[0033] In the formula, p is the real-time complete data point, N k (p) is the local data group corresponding to the real-time complete data point p, q is the data point in the local data group corresponding to the real-time complete data point, MD(q) is the density of the data point q; MD(p) is the density of the real-time complete data point p, K is the number of neighbors, is the local density between the real-time complete data point p and the data point q.
[0034] Preferably, the specific steps of S4 also include:
[0035] S43. Based on the value of the anomaly index YZ(p) at the real-time complete data point obtained in S42, identify the anomaly points, and the specific steps are as follows:
[0036] S431. If the value of the anomaly index YZ(p) at the real-time complete data point = 1, it indicates that the density of the real-time complete data point p is similar to that of its neighbors, and it is determined that the real-time complete data point p is a normal point in the local data group;
[0037] S432. If the value of the anomaly index YZ(p) at the real-time complete data point > 1, it indicates that the density of the real-time complete data point p is lower than that of the neighbor points, and it is considered that the real-time complete data point p is an anomaly point. At this time, an anomaly notification will be sent to the user end;
[0038] S433. If the value of the anomaly index YZ(p) at the real-time complete data point < 1, it indicates that the density of the real-time complete data point p is higher than that of the neighbor points, and it is determined that the real-time complete data point p is a normal point within the local data group.
[0039] Preferably, the specific steps of S4 further include:
[0040] S44. According to the anomaly notification and combined with the anomaly index YZ(p) at the real-time complete data point, dynamically adjust the number of neighbors K, specifically:
[0041]
[0042] In the formula, K new is the adjusted number of neighbors, α is the weight coefficient, α ∈ (0, 1); YZ(p) σ is the standard deviation of the anomaly index at the real-time complete data point; is the density uniformity index.
[0043] A user-side abnormal power consumption monitoring system based on data analysis includes a preparation subsystem, a test subsystem, a compensation subsystem, and an anomaly recognition and optimization subsystem;
[0044] The preparation subsystem is used to monitor the real-time power consumption data of the user side through the sensor group, analyze the periodic dependence degree of the sequence in combination with the timing, so as to set the window upper limit T max at each data point, and based on the historical data group, screen the windows without missing data from the window upper limit T max at each data point and mark them as complete windows;
[0045] The test subsystem is used to set different missing ranges according to the complete windows to perform the missing compensation means test operation, and select the corresponding missing compensation means according to different missing conditions to construct a compensation set;
[0046] The compensation subsystem is used to preprocess the real-time power consumption data, and combine the window upper limit T max at the data point obtained in S1 and S2 and the compensation set to take the missing compensation means for the real-time data point to generate real-time complete data points;
[0047] The anomaly recognition and optimization subsystem is used to preset the number of neighbors K in advance, determine the K neighbor points adjacent to the real-time complete data point, and analyze and compare the density situation between the real-time complete data point and its neighbor points to judge the anomaly points, so as to monitor the abnormal power consumption of the user side and dynamically adjust the number of neighbors K.
[0048] The present invention provides a user-side abnormal power consumption monitoring method and system based on data analysis, which has the following beneficial effects:
[0049] (1) By analyzing the periodic dependence degree of the sequence according to the time series, setting the window upper limit at each data point, and screening out the complete windows without missing values, the influence of data missing on the monitoring results can be effectively eliminated, ensuring the reliability of subsequent data analysis. Especially in electricity consumption data, improper handling of missing values may lead to monitoring errors, while the complete window selection method proposed by the present invention can ensure the integrity of data, providing accurate data support for abnormal electricity consumption detection. Optimize the missing value compensation strategy: According to the different missing ranges of historical data and real-time data, combined with different missing conditions, appropriate missing value compensation means are selected, and a compensation set is constructed, avoiding the singularity and limitations of traditional missing value filling methods. By dynamically adapting to different missing scenarios, the present invention can fill in missing data more accurately, reduce data deviation, and further improve the effectiveness of data. Enhance the abnormal detection ability: By presetting and dynamically adjusting the number of neighbors K, and using the density analysis method to judge whether the real-time data point is abnormal, it can flexibly respond to the changes in the electricity consumption patterns of different user terminals and identify abnormal electricity consumption behaviors in real time. The practice of dynamically adjusting the K value can automatically optimize the sensitivity of abnormal detection according to the change of data density, reducing the possibility of false positives and false negatives. Improve the system self-adaptability and flexibility: The method of the present invention can automatically adjust the processing strategy according to the characteristics and missing situations of different electricity consumption data, improving the adaptability and flexibility of the system. Especially in the face of the dynamic changes of real-time electricity consumption data, it can effectively adjust the parameters of data processing and abnormal monitoring algorithms, ensuring the stability and reliability of the system in a changing environment. In summary, through the combination of time series analysis, compensation means optimization and density analysis, this method can effectively improve the accuracy of abnormal electricity consumption monitoring, and adaptively adjust parameters during the processing, fully solving the problems of inaccurate missing value filling and unstable abnormal detection in traditional methods, and is applicable to various complex electricity consumption monitoring scenarios.
[0050] (2)Precisely capture the dependencies and periodic changes between data points in the time series based on historical data groups, which can effectively identify the correlations between different time points. Especially in the dynamic user-side electricity consumption scenarios, it can accurately evaluate the change trends of electricity consumption data under different lag steps, providing a scientific basis for subsequent anomaly monitoring and prediction. Dynamic adjustment of the window upper limit and adaptive analysis: In step S13, by plotting the relationship graph based on the lag correlation coefficient and analyzing the change trends of different lag steps, the window upper limit of each data point can be flexibly set according to the actual changes in the data, avoiding the monitoring errors that may be caused by a fixed window size. The maximum drop difference method ensures sensitive capture of data changes, making the setting of the window upper limit more accurate, thus effectively optimizing the data point screening and anomaly identification process. Data integrity guarantee and precise screening: In step S14, extract and screen the complete windows without missing data from the historical data group to ensure the integrity and accuracy of the data set used in the monitoring process. By this method, the impact of data missing on the monitoring results can be further reduced, avoiding anomaly detection errors caused by incomplete data.
[0051] (3)Through calculating the anomaly index, the present invention can dynamically determine whether a data point is abnormal according to the density comparison between the real-time complete data point and its neighbor points. For data points with low density, the system will issue an alarm and notify the user side in a timely manner. This method does not require pre-setting a fixed anomaly threshold, but adaptively determines anomalies according to the local characteristics of the data, improving the flexibility and accuracy of the monitoring strategy. In short, through accurately identifying anomaly points, this method can help power companies or equipment operators efficiently manage electricity resources and early warn potential equipment failures or electricity resource abuse phenomena. When deployed on a large scale, this system can effectively reduce system risks and ensure power supply safety and reliability.
[0052] (4)Through the dynamic adjustment mechanism of the present invention, the detection of anomaly points no longer depends on fixed parameters, but can be automatically optimized according to the real-time changes in the data. This not only improves the accuracy of anomaly monitoring, but also reduces the complexity of manual parameter adjustment, making the system more intelligent and having the ability of autonomous learning and adaptation. In summary, the method of dynamically adjusting the nearest neighbor number K based on the anomaly index can effectively optimize the accuracy and flexibility of anomaly detection, enhance the adaptability of the system in different data density scenarios, make the anomaly monitoring of user-side electricity consumption more intelligent, efficient and stable, while reducing the computational complexity and resource consumption, and comprehensively improving the overall performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a schematic flow chart of a method for monitoring abnormal electricity consumption of user side based on data analysis according to the present invention;
[0054] Figure 2It is the relational diagram involved in S13 in the present invention;
[0055] Figure 3 It is the block diagram of an abnormal power consumption monitoring system for the client based on data analysis in the present invention; Specific implementation manner
[0056] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0057] Embodiment 1
[0058] Please refer to Figure 1 , the present invention provides an abnormal power consumption monitoring method for the client based on data analysis, including the following steps:
[0059] S1: Monitor the real-time power consumption data of the client through the sensor group, combine the timeliness, analyze the periodic dependence degree of the sequence, so as to set the window upper limit T at each data point max , and based on the historical data group, screen the windows without missing data from the window upper limit T at each data point max , and mark them as complete windows;
[0060] S2: According to the complete window, set different missing ranges to perform the missing compensation means test operation, and select the corresponding missing compensation means according to different missing conditions to construct a compensation set;
[0061] S3: After preprocessing the real-time power consumption data, combine the window upper limit T at the data points obtained in S1 and S2 max and the compensation set, and take the missing compensation means for the real-time data points to generate real-time complete data points;
[0062] S4: Preset the number of neighbors K in advance, determine the K neighbor points adjacent to the real-time complete data points, and analyze and compare the density between the real-time complete data points and their neighbor points to judge the abnormal points, so as to monitor the abnormal power consumption of the client and dynamically adjust the number of neighbors K.
[0063] In this embodiment, by combining temporal and periodic dependence analysis, complete data is screened within a set window, avoiding the errors and biases brought by traditional random missing value filling methods. This screening mechanism based on the window upper limit ensures data integrity, and thus provides a more reliable basis for subsequent abnormal power consumption monitoring. Flexible missing compensation strategy: According to the data within the complete window, different missing ranges are set and compensation means are tested, enabling dynamic selection of relatively optimal missing compensation means for different missing conditions. The constructed compensation set provides a more flexible and accurate processing method for various missing scenarios, effectively improving the accuracy in the data preprocessing stage. Effective abnormal power consumption monitoring: After preprocessing the real-time power consumption data, by combining the window upper limit and the compensation set for missing compensation, high-quality real-time complete data points can be generated. This process ensures the improvement of data quality and lays a good foundation for subsequent abnormal power consumption monitoring. Adaptive neighbor analysis mechanism: By setting the number of neighbors K and dynamically adjusting this parameter, the density analysis of neighbor points can be adjusted according to the actual situation of the data, thereby improving the accuracy of abnormal point detection. Dynamically adjusting the value of K enables this method to flexibly adapt to the density change of the data, optimizing the abnormal detection result and reducing false alarms and missed detections. Improving monitoring efficiency and accuracy: Considering temporal dependence, historical data, missing compensation means, and neighbor point density analysis comprehensively, the present invention provides an all-round and refined abnormal monitoring strategy. This method can effectively improve the recognition ability of the monitoring system for abnormal power consumption behaviors, optimize the overall efficiency and accuracy of power consumption monitoring, and reduce the waste of system resources and unnecessary manual intervention. In summary, the present invention can achieve accurate identification and monitoring of abnormal power consumption situations at the user end, improve the intelligent level of power consumption data analysis, and has strong practical value and application prospects.
[0064] Embodiment 2
[0065] Please refer to Figure 1 and Figure 2 , specifically: The specific steps of S1 include:
[0066] S11. Use a sensor group to monitor the real-time power consumption data of the user end. Among them, the real-time power consumption data includes, but is not limited to, the power consumption, voltage, current, power factor, electric energy, load, timestamp, temperature, humidity, and grid frequency of the user end at the real-time data point. Based on the real-time power consumption data, an electricity consumption feature vector is generated; and the real-time user data is regularly stored in a cloud database according to the time series to generate a historical data group, and the historical data group includes, but is not limited to, the power consumption, voltage, current, power factor, electric energy, load, timestamp, temperature, humidity, and grid frequency of the user end at the historical data point;
[0067] S12. Analyze the periodic dependence among time series data based on the historical data group to obtain the lag correlation coefficient Z(k). Specifically, it is obtained through the following method:
[0068]
[0069] In the formula, Z(k) is the lag correlation coefficient, representing the lag correlation coefficient of the time series at lag k; X t is the electricity consumption feature vector at data point t; X avg is the mean value of the electricity consumption feature vector; X t+k is the electricity consumption feature vector at data point t + k; T is the length of historical data points in the historical data group; t is the number of the data point; k is the lag step. Among them, using T - k as the summation range is to ensure that when calculating the lag correlation coefficient Z(k), a lag of k will not cause the data to exceed the actual length of the time series.
[0070] Specifically, in S12, the system analyzes the periodic dependence between time series data based on the historical data group. The periodic dependence in the time series refers to the mutual relationship between a data point and its previous and subsequent time points. These relationships may reflect factors such as the electricity consumption pattern of the device and seasonal changes. The lag correlation coefficient is an important indicator describing the degree of correlation of time series data at different lag steps. In the calculation formula of the lag correlation coefficient, the involved electricity consumption feature vector and its mean value are used to quantify the correlation between data points: by calculating the correlation between these two time points, the lag effect of the data can be revealed. For example, a larger correlation coefficient for lag step k indicates a stronger correlation between the current electricity consumption state and the electricity consumption state k steps ago.
[0071] The specific steps of S1 also include:
[0072] S13. Based on the lag correlation coefficient Z(k) obtained in S12, draw a relationship graph, and based on the relationship graph, obtain the change trend of the time series at different lag steps k, and based on the change trend of the time series at different lag steps k, set the window upper limit T at each data point max , and the specific setting steps are as follows:
[0073] S131. Based on the relationship graph, determine the decline rate Jz between each data point and the data points at different lag steps k, and compare the adjacent two decline rates Jz in the relationship graph to obtain the decline difference Jc;
[0074] S132. Through statistics, determine the maximum decline difference Jc max , and use the position corresponding to the maximum decline difference Jc max in the relationship graph as the upper limit of the lag point, and based on the upper limit of the lag point, determine the window upper limit T at the corresponding data pointmax 。
[0075] Among them, the relationship graph shows the change trend of the time series at different lag steps. By observing the relationship graph, the change law of the data can be identified, and then the upper limit of the window for each data point can be set. The upper limit of the window refers to the maximum range of data change when a given lag step is set.
[0076] Among them, the decline rate Jz represents the change amplitude between data points as the lag step k changes. The calculation of the decline rate can reveal the degree of change of electricity consumption data under the influence of lag.
[0077] Among them, the decline difference Jc is the difference between the decline rates of adjacent data points. Through the statistical analysis of the decline difference, the system can identify the maximum decline difference and set the upper limit of the lag point accordingly, that is, determine the change trend of the data points at different lag steps.
[0078] The specific steps of S1 also include:
[0079] S14. Extract several upper limits T of the windows at the data points from the historical data group max and screen the windows without missing data from the upper limits T of the windows at each data point max to be marked as complete windows.
[0080] Each data point will have an upper limit according to the lag analysis. For each data point, if the data in its corresponding window has no missing or abnormal data, it is marked as a complete window. The data of these complete windows will be used for further analysis, modeling and prediction.
[0081] In this embodiment, by monitoring in real time and storing power consumption data in the cloud database at regular intervals, historical data groups can be generated in real time, and long-term tracking of the power consumption behavior of the user side can be achieved through continuous storage. The application of the cloud database makes data storage more secure and scalable, and at the same time provides rich historical data support for subsequent analysis, which helps to deeply analyze power consumption patterns and behavioral characteristics. By introducing the calculation of the lag correlation coefficient and combining the power consumption feature vectors in the historical data group for periodic dependence analysis, the periodic changes and interdependent relationships in the time series can be accurately captured. This analysis method can reveal the correlation coefficients between different time points, and then accurately evaluate the time dependence between data points, providing a theoretical basis for subsequent window setting and data processing. Flexible window upper limit setting: By plotting the lag correlation coefficient relationship diagram and analyzing the change trend at different lag steps k, the window upper limit at each data point can be accurately determined. Based on this setting, the window range of data points can be adjusted more flexibly, so as to more effectively segment and fill missing values of the data. In addition, the dynamic setting of the window upper limit according to the maximum drop difference Jc of the lag point upper limit can ensure that information loss or analysis deviation will not be caused by too long or too short window ranges during the data processing and analysis process. Complete window screening and missing data elimination: By screening the complete windows in the historical data group, that is, the windows without missing data, the integrity of the data used in the subsequent analysis process is guaranteed. Marking and screening out the complete windows without missing data can ensure that the data processing stage is not affected by missing data, and thus avoid the interference of data missing on the effectiveness of the anomaly detection algorithm, providing high-quality data support for abnormal power consumption monitoring. To sum up, the monitoring method of the present invention significantly improves the accuracy of data analysis and the self-adaptability of the system through steps such as accurate time series analysis, window upper limit setting, lag correlation coefficient calculation, and complete window screening. This method can efficiently handle the missing problems in power consumption data, optimize the accuracy and stability of abnormal power consumption monitoring, is applicable to various user-side power consumption monitoring scenarios, and has a wide application prospect.
[0082] Embodiment 3
[0083] Please refer to Figure 1 , specifically: The specific steps of S2 include:
[0084] S21. Set different missing ranges for the complete window. The missing range refers to the range where missing values exist within the window corresponding to each data point, and according to different missing conditions, conduct missing compensation means test operations on the missing values within the corresponding data points. Among them, the missing compensation means include but are not limited to mean filling, median filling, mode filling, forward filling, backward filling, and interpolation method;
[0085] Specifically, the missing value compensation method: There are various methods for filling missing values, including: Mean filling: Fill the missing value with the mean of other data points within the window. Median filling: Fill the missing value with the median of other data points within the window. Mode filling: Fill the missing value with the most frequently occurring value within the window. Forward filling: Fill the missing value with the value at the previous time point. Backward filling: Fill the missing value with the value at the next time point. Interpolation method: Estimate the missing value through an interpolation algorithm (such as linear interpolation).
[0086] S22. During the process of performing the missing value compensation method test operation, determine the filling results of the missing values by various missing value compensation methods under different missing conditions to obtain the filled value BT. Based on the filled value BT, identify the corresponding missing value compensation methods under different missing conditions. The specific identification steps are as follows:
[0087] S221. Set different missing conditions, specifically: where Tcb t is the missing value ratio corresponding to the data point t; Qc t is the missing value range within the window upper limit T max at the data point t; T max,t is the window upper limit T max at the data point t; the missing value ratio Tcb t at the data point t reflects the set missing conditions at the corresponding data point;
[0088] S222. Obtain the deviation value PT, and the deviation value PT is obtained through the following formula: PT = |BT - SZ|. The calculation method in the formula is used to quantify the effects of different compensation methods.
[0089] where SZ is the actual value; the deviation situation of filling the missing values by various missing value compensation methods under the corresponding missing conditions is reflected through the deviation value PT;
[0090] The deviation value reflects the difference between the filled data and the real data. The smaller the deviation value, the better the filling effect and the more accurate the filling method.
[0091] S223. After statistics, select the minimum deviation value PT min , and use the missing value compensation method corresponding to the minimum deviation value PT min as the best missing value compensation method for the set missing conditions within the complete window at the corresponding data point to construct a compensation set. The compensation set includes the best missing value compensation methods for different missing conditions set within the complete window at the corresponding data point.
[0092] In this embodiment, by setting different missing ranges and testing compensation means for each data point according to different missing conditions, the relatively appropriate compensation means required under different missing conditions can be accurately identified. The diverse missing compensation methods (such as mean filling, forward filling, etc.) ensure that in different data missing scenarios, the relatively optimal compensation means can be automatically selected, avoiding the errors that may be brought by random compensation methods, thus guaranteeing the accuracy and integrity of the data. Reducing the impact of data missing on the results: By accurately evaluating the deviation values of each missing compensation means and selecting the compensation method corresponding to the minimum deviation value, the adverse impact of data missing on subsequent algorithm analysis (such as anomaly detection, trend prediction, etc.) can be minimized as much as possible. The optimization of the compensation method enables the monitoring system to better maintain the stability of the original data trend when facing missing data, thereby enhancing the robustness of the entire abnormal power consumption monitoring process.
[0093] Dynamic adaptation to different missing conditions: By setting missing conditions according to the missing value ratio at the data point and selecting corresponding compensation means for different missing situations, the system can be flexibly adapted to different data missing types. In practical applications, the power consumption data missing situations in different time periods or different user terminals may be different. The method of the present invention can intelligently identify and effectively cope with these differences, enabling the monitoring system to maintain high stability and accuracy in various scenarios. Optimization of compensation set construction: By testing and comparing various compensation means under different missing conditions, selecting the compensation means corresponding to the minimum deviation value as the best filling strategy, and constructing a compensation set based on this strategy, the entire data processing flow is further optimized. The construction of the compensation set provides a more accurate compensation basis for subsequent data processing, thereby improving the reliability of subsequent data analysis and anomaly detection. Improving the accuracy and sensitivity of anomaly monitoring: Due to the optimization of the missing value compensation means, the integrity of the data is effectively guaranteed, enabling the subsequent abnormal power consumption monitoring process to be analyzed based on more accurate data, reducing the risk of false alarms or missed abnormal phenomena caused by data missing. The dynamic selection of the compensation set further enhances the ability of the monitoring system to capture minor abnormal fluctuations, ensuring the efficient and accurate identification of abnormal power consumption behaviors by the system during real-time monitoring. In summary, through the accurate analysis and dynamic compensation of missing data, the present invention can not only improve the integrity and accuracy of the data, but also flexibly adjust the compensation strategy under different missing conditions, thereby effectively enhancing the reliability and practicality of abnormal power consumption monitoring at the user end.
[0094] Embodiment 4
[0095] Please refer to Figure 1 , specifically: The specific steps of S3 include:
[0096] S31. Preprocess the real-time electricity consumption data obtained in S1 to identify the missing values in the real-time electricity consumption data, determine the range of missing values, and combine the window upper limit T at the data points obtained in S1 and S2 max and the compensation set, and adopt corresponding optimal missing compensation means for the data points corresponding to the real-time electricity consumption data to generate real-time complete data points.
[0097] Missing values are sometimes represented by null values or empty strings. In this case, the null values or empty strings in the data can be directly checked, or they can be identified by visualization methods.
[0098] In this embodiment, by preprocessing the real-time electricity consumption data, the missing values in the data can be timely identified and their missing ranges can be determined. This accurate identification of missing values ensures the integrity of the data and provides high-quality input data for subsequent anomaly monitoring. Combining the window upper limit and the compensation set obtained in S1 and S2, a relatively suitable compensation method can be automatically selected and applied for each data point, making the data recovery more accurate and avoiding the impact of improper filling on the monitoring results. The present invention dynamically adjusts and selects the optimal missing compensation means, and implements a tailored compensation strategy according to the missing situation of the real-time electricity consumption data to ensure that the filling method for each data point is relatively appropriate.
[0099] Embodiment 5
[0100] Please refer to Figure 1 , specifically: The specific steps of S4 include:
[0101] S41. Preset the number of neighbors K, respectively determine the Euclidean distances between each historical data point in the historical data group and the real-time complete data point, and extract K groups of historical data points adjacent to the real-time complete data point from them to obtain a local data group;
[0102] S42. Based on the local data group, analyze and compare the density between the real-time complete data point and its neighbor points to obtain the anomaly index YZ(p) at the real-time complete data point, which is specifically obtained through the following formula:
[0103]
[0104] In the formula, p is the real-time complete data point (the electricity consumption situation of the user side at the current moment), N k (p) is the local data group corresponding to the real-time complete data point p, q is the data point in the local data group corresponding to the real-time complete data point, MDd(q) is the density of the data point q; MD(p) is the density of the real-time complete data point p, and K is the number of neighbors. Let $\rho(p,q)$ be the local density of the real-time complete data point $p$ and the data point $q$. The larger the value, the lower the density of $p$ relative to $q$, that is, $p$ is more likely to be an outlier.
[0105] Among them, the density of a data point refers to the relative "crowdedness" of the data point within its neighborhood. The higher the density, the more similar data points there are in the area around the data point. The lower the density, the sparser the area around the data point. Specifically, it can be obtained through the following methods:
[0106]
[0107] Among them, $d(q,p$ i ) represents the Euclidean distance between the data point $p$ and its $j$-th nearest neighbor data point $q$ i ; $j$ represents the number of the nearest neighbor data point;
[0108] The specific steps of S4 also include:
[0109] S43. Based on the value of the outlier index $YZ(p)$ at the real-time complete data point obtained in S42, identify the outlier. The specific steps are as follows:
[0110] S431. If the value of the outlier index $YZ(p)$ at the real-time complete data point = 1, it indicates that the density of the real-time complete data point $p$ is similar to that of its neighbors, and it is determined that the real-time complete data point $p$ is a normal point within the local data group;
[0111] S432. If the value of the outlier index $YZ(p)$ at the real-time complete data point > 1, it indicates that the density of the real-time complete data point $p$ is lower than that of its neighbor points, and the density of the real-time complete data point $p$ is significantly lower than that of its neighbors. It is considered that the real-time complete data point $p$ is an outlier. At this time, an outlier notification will be sent to the user side;
[0112] S433. If the value of the outlier index $YZ(p)$ at the real-time complete data point < 1, it indicates that the density of the real-time complete data point $p$ is higher than that of its neighbor points, and it is determined that the real-time complete data point $p$ is a normal point within the local data group.
[0113] In this embodiment, by presetting the number of neighbors K and calculating the Euclidean distances between each historical data point in the historical data group and the real-time complete data point, the K neighbor points adjacent to the real-time data point can be efficiently determined. The construction of this local data group enables the anomaly detection of data points to be analyzed based on the local environment, enhancing the accuracy of detection and avoiding misjudgments that may be caused by global data fluctuations. Precise determination of the anomaly index: In step S42, by calculating the anomaly index at the real-time complete data point based on density differences and judging whether the data point is abnormal through the value of the anomaly index, the abnormal electricity consumption behavior of the user side can be more precisely identified. Specifically, when the value of the anomaly index is significantly greater than 1, it indicates that the density of the real-time data point is lower than that of the neighbor data points, thus indicating that this data point is likely to be an outlier. This judgment method based on density differences is more flexible than the traditional threshold judgment and can be automatically adjusted according to the changes in real-time data. Dynamic adjustment and real-time response: Based on the calculation and comparison of the anomaly index, the system can dynamically identify the anomaly points with significantly different densities from the real-time complete data point and send an anomaly notification to the user side in a timely manner. This real-time response mechanism ensures that the system can quickly detect abnormal electricity consumption situations, provide timely feedback and warnings to users, and effectively avoid the risks brought by potential power abuse or equipment failures. By setting the number of neighbors K and combining the analysis method of local density, the method of the present invention has good scalability. The dynamic adjustment of the K value enables this method to adapt to different types of user-side data and can also be flexibly applied to large-scale user-side data, providing higher flexibility and versatility.
[0114] Embodiment 6
[0115] Please refer to Figure 1 , specifically: The specific steps of S4 further include:
[0116] S44. According to the anomaly notification and in combination with the anomaly index YZ(p) at the real-time complete data point, dynamically adjust the number of neighbors K, specifically:
[0117]
[0118] In the formula, K new is the adjusted number of neighbors, α is the weight coefficient, which controls the influence degree of density uniformity on the adjustment of the number of neighbors K, and α ∈ (0, 1); YZ(p) σ is the standard deviation of the anomaly index at the real-time complete data point. The smaller the value, the higher the density uniformity; is the density uniformity index. The larger the value, the more uniform the data set. The adjusted number of neighbors K new will be larger. If the density fluctuates greatly (the standard deviation YZ(p) of the anomaly index at the real-time complete data point σ ), then the density uniformity index is smaller, which means that the adjusted number of neighbors Knew The smaller the adjustment range.
[0119] For sparse or non-uniform data sets, a smaller K value can better capture the changes between sparse points or dense points in the data. When the density of the data set is relatively uniform, the distribution of data points is more consistent, without prominent sparse areas or aggregation areas. In this case, a larger K value can help the algorithm better smooth the changes in local density.
[0120] In this embodiment, the adaptive adjustment of the number of neighbors: The present invention dynamically adjusts the number of neighbors K according to the standard deviation of the anomaly index and the density uniformity index, enabling the algorithm to flexibly cope with the density characteristics of different data sets. For sparse or non-uniform data sets, a smaller K value can help the algorithm capture the changes between sparse points or dense points in the data, enhancing the sensitivity of anomaly detection; while when the data set density is uniform, a larger K value helps the algorithm smooth the local density changes and improve the stability of anomaly detection. This adaptive adjustment mechanism ensures that the algorithm can optimize the anomaly detection results according to the characteristics of real-time data, avoiding the limitations brought by a fixed K value. By adjusting the number of neighbors K, it can accurately adapt to the distribution of data, avoid over-smoothing or over-refining the changes in local density, and further improve the accuracy of anomaly detection. For data sets with large density fluctuations, appropriately reducing the adjustment range of the K value helps avoid over-amplifying the density differences of anomaly points, thereby reducing false alarms; while for data sets with uniform density, appropriately increasing the K value can improve the accuracy and stability of detection and effectively prevent anomaly points from being ignored.
[0121] Optimizing resource consumption and computing efficiency: During the process of dynamically adjusting the K value, a smaller K value will reduce the consumption of computing resources, especially for sparse or non-uniformly distributed data sets, which helps reduce the computational complexity and improve the computing efficiency. At the same time, a larger K value helps smooth the local density changes when the data set density is high, avoiding overfitting, thereby effectively saving computing resources and improving the system's response speed. Enhancing the robustness and adaptability of the algorithm: By introducing the density uniformity index and the standard deviation of the anomaly index to adjust the number of neighbors K, the algorithm can more accurately judge the characteristics of the data set and make corresponding adjustments. This method enables the system to handle variable data in different environments, such as fluctuations in electricity consumption data, seasonal changes, etc., enhancing the robustness and adaptability of the algorithm in practical applications and ensuring the effectiveness of anomaly detection.
[0122] Embodiment 7
[0123] Please refer to Figure 3 , specifically: A user-side abnormal electricity consumption monitoring system based on data analysis, including a preparation subsystem, a test subsystem, a compensation subsystem, and an anomaly identification and optimization subsystem;
[0124] A preparation subsystem for monitoring real-time power consumption data of a user terminal through a sensor group, analyzing the periodic dependence degree of a sequence in combination with time series, so as to set an upper window limit T at each data point max , and based on a historical data group, screening windows without missing data from the upper window limit T at each data point max and marking them as complete windows;
[0125] A testing subsystem for setting different missing ranges according to the complete windows to perform a missing compensation means testing operation, and selecting corresponding missing compensation means according to different missing conditions to construct a compensation set;
[0126] A compensation subsystem for, after preprocessing the real-time power consumption data, taking a missing compensation means for real-time data points in combination with the upper window limit T at the data points obtained in S1 and S2 max and the compensation set to generate real-time complete data points;
[0127] An anomaly recognition and optimization subsystem for presetting a number of neighbors K, determining K neighbor points adjacent to the real-time complete data points, and analyzing and comparing the density between the real-time complete data points and their neighbor points to judge anomaly points, so as to monitor abnormal power consumption of the user terminal and dynamically adjust the number of neighbors K.
[0128] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for monitoring abnormal power consumption at a user end based on data analysis, characterized in that: The following steps are involved: S1: Monitor the real-time power consumption data of the user end through the sensor group, analyze the periodic dependence of the sequence in combination with the timing, and set the upper limit T of the window at each data point max , and based on the historical data set, the upper limit T of the window at each data point max The windows without missing data are selected internally and marked as complete windows; S2: according to the complete window, different missing ranges are set to perform missing compensation means testing operations, and corresponding missing compensation means are selected according to different missing conditions to construct a compensation set; S3: The real-time power consumption data is preprocessed and combined with the window upper limit T at the data point obtained in S1 and S2 max and compensation set, taking missing compensation measures for real-time data points to generate real-time complete data points; S4: Pre-set the number of neighbors K, determine the K neighboring points adjacent to the real-time complete data point, analyze and compare the density between the real-time complete data point and its neighboring points, judge the abnormal points, monitor the abnormal power consumption at the user end, and dynamically adjust the number of neighbors K.
2. According to claim 1, a method for monitoring abnormal power consumption at a user end based on data analysis is characterized in that: The specific steps of S1 include: S11, using a sensor group to monitor the real-time power consumption data of the user end, generating a power consumption feature vector based on the real-time power consumption data; and storing the real-time user data in a cloud database in a time series to generate a historical data group; S12. Based on the historical data group, analyze the periodic dependence between the time series data to obtain the lagged correlation coefficient Z(k), which is obtained specifically in the following way: Where, X t is the electricity consumption feature vector at data point t; X avg is the mean of the electricity consumption characteristic vector; X t+k is the electricity consumption characteristic vector at data point t+k; T is the length of the historical data points in the historical data group; t is the number of the data point; k is the number of lag steps.
3. The method for monitoring abnormal power consumption at a user end based on data analysis according to claim 2 is characterized in that: The specific steps of S1 also include: S13, based on the lag correlation coefficient Z(k) obtained in S12, draw a relationship diagram, and according to the relationship diagram, obtain the change trend of the time series at different lag steps k, and set the window upper limit T at each data point based on the change trend of the time series at different lag steps k max , the specific setting steps are as follows: S131, based on the relationship graph, determine the drop rate Jz between each data point and data points with different lag steps k, and compare two adjacent groups of drop rates Jz in the relationship graph to obtain the drop difference Jc; S132, through statistics, determine the maximum drop difference Jc max , and the maximum drop difference Jc max The corresponding position in the relationship diagram is used as the upper limit of the hysteresis point. According to the upper limit of the hysteresis point, the upper limit of the window T at the corresponding data point is determined. max .
4. The method for monitoring abnormal power consumption at a user end based on data analysis according to claim 3 is characterized in that: The specific steps of S1 also include: S14, extracting the upper limit T of the window at several groups of data points from the historical data group max , and the upper limit T of the window at each data point max Filter windows without missing data to mark them as complete windows.
5. The method for monitoring abnormal power consumption at a user end based on data analysis according to claim 4 is characterized in that: The specific steps of S2 include: S21. Setting different missing ranges for the complete window, where the missing range refers to the range of missing values in the window corresponding to each data point, and performing missing compensation means testing operations on the missing values in the corresponding data points according to different missing conditions; S22. In the process of performing the missing compensation means test operation, the filling results of the missing values by various missing compensation means under different missing conditions are determined to obtain the filling value BT. Based on the filling value BT, the corresponding missing compensation means under different missing conditions are identified. The specific identification steps are: S221. Set different missing conditions, specifically: Among them, Tcb t is the missing value ratio corresponding to data point t; Qc t is the upper limit of the window at data point t max The missing value range within T max,t is the upper limit of the window at data point t max ; S222, obtaining a deviation value PT, wherein the deviation value PT is obtained by the following formula: PT=|BT-SZ|, wherein SZ is an actual value; S223, after statistics, select the minimum deviation value PT min , and the minimum deviation value PT min The corresponding missing compensation means is used as the optimal missing compensation means for the missing conditions set within the complete window at the corresponding data point to construct a compensation set, which includes the optimal missing compensation means for different missing conditions set within the complete window at the corresponding data point.
6. A method for monitoring abnormal power consumption at a user end based on data analysis according to claim 5, characterized in that: The specific steps of S3 include: S31, preprocessing the real-time power consumption data obtained in S1 to identify missing values in the real-time power consumption data and determine the range of missing values, combining the window upper limit T at the data points obtained in S1 and S2 max and compensation sets, taking corresponding optimal missing compensation means for the data points corresponding to the real-time electricity consumption data to generate real-time complete data points.
7. A method for monitoring abnormal power consumption at a user end based on data analysis according to claim 6, characterized in that: The specific steps of S4 include: S41, presetting the number of nearest neighbors K, determining the Euclidean distance between each historical data point and the real-time complete data point in the historical data group, and extracting K groups of historical data points adjacent to the real-time complete data point to obtain a local data group; S42, based on the local data group, analyzing and comparing the density between the real-time complete data point and its neighboring points to obtain the abnormal index YZ(p) at the real-time complete data point, which is specifically obtained by the following formula: Where p is the real-time complete data point, N k (p) is the local data group corresponding to the real-time complete data point p, q is the data point in the local data group corresponding to the real-time complete data point, MD(q) is the density of data point q; MD(p) is the density of real-time complete data point p, K is the number of neighbors, is the local density of the real-time complete data point p and the data point q.
8. The method for monitoring abnormal power consumption at a user end based on data analysis according to claim 7 is characterized in that: The specific steps of S4 also include: S43, based on the abnormal index YZ(p) value at the real-time complete data point obtained in S42, identifying the abnormal point, the specific steps are as follows: S431, if the value of the abnormal index YZ(p) at the real-time complete data point is 1, it indicates that the density of the real-time complete data point p is similar to the density of its neighbors, and the real-time complete data point p is determined to be a normal point in the local data group; S432, if the value of the abnormal index YZ(p) at the real-time complete data point is greater than 1, it indicates that the density of the real-time complete data point p is lower than that of the neighboring points, and the real-time complete data point p is considered to be an abnormal point, and an abnormal notification is sent to the user end; S433. If the value of the abnormal index YZ(p) at the real-time complete data point is less than 1, it indicates that the density of the real-time complete data point p is higher than that of its neighboring points, and the real-time complete data point p is determined to be a normal point in the local data group.
9. A method for monitoring abnormal power consumption at a user end based on data analysis according to claim 8, characterized in that: The specific steps of S4 also include: S44. According to the abnormal notification and in combination with the abnormal index YZ(p) at the real-time complete data point, the number of neighbors K is dynamically adjusted, specifically: In the formula, K new is the adjusted number of neighbors, α is the weight coefficient, α∈(0,1); YZ(p) σ is the standard deviation of the anomaly index at the real-time complete data points; It is an indicator of density uniformity.
10. A user-side abnormal power consumption monitoring system based on data analysis, used to implement a user-side abnormal power consumption monitoring method based on data analysis as claimed in any one of claims 1 to 9, characterized in that: It includes preparation subsystem, test subsystem, compensation subsystem, abnormality identification and optimization subsystem; Prepare a subsystem to monitor the real-time power consumption data of the user end through the sensor group, analyze the periodic dependence of the sequence in combination with the timing, and set the upper limit T of the window at each data point max , and based on the historical data set, the upper limit T of the window at each data point max The windows without missing data are selected internally and marked as complete windows; A test subsystem, used for setting different missing ranges according to the complete window to perform missing compensation means testing operations, and selecting corresponding missing compensation means according to different missing conditions to construct a compensation set; The compensation subsystem is used to combine the real-time power consumption data with the window upper limit T at the data point obtained in S1 and S2 after preprocessing. max and compensation set, taking missing compensation measures for real-time data points to generate real-time complete data points; The anomaly identification and optimization subsystem is used to pre-set the number of neighbors K, determine the K neighboring points adjacent to the real-time complete data point, analyze and compare the density between the real-time complete data point and its neighboring points, judge the abnormal points, monitor the abnormal power consumption at the user end, and dynamically adjust the number of neighbors K.
Citation Information
Patent Citations
Power data abnormal value detection algorithm based on K neighbor density peak value clustering
CN114417971A
Complementation method and system using data collected by transformer area
CN115439102A
Thermodynamic data anomaly detection and restoration method based on mechanism and data cooperative driving
CN117574290A
Energy management system data filling method and architecture based on time sequence analysis
CN119377203A
Method for processing electricity consumption data and server implementing same
WO2024122786A1