Intelligent power consumption behavior inspection method and device based on multi-source data and storage medium
Through distributed computing and deep learning technology, combined with sliding window algorithms and correlation rule mining, the problems of inefficiency and insufficient accuracy of traditional power consumption auditing methods are solved, intelligent analysis and real-time monitoring of power consumption behavior are realized, and the accuracy and efficiency of auditing are improved.
Patent Information
- Application Number
- CN202510704843.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
Traditional electricity inspection methods are inefficient, easily affected by human factors, difficult to monitor massive data in real time, and lack the ability to integrate multi-dimensional information, resulting in insufficient accuracy in detecting abnormal electricity use behaviors.
The distributed computing framework is used to process the power data flow, combine the sliding window algorithm to extract real-time changing features, and mining through multi-dimensional feature vectors and association rules, using deep learning technology to confirm abnormal behavior, and using visualization technology to display the audit results.
It realizes intelligent analysis and real-time monitoring of electricity consumption behavior, improves the accuracy and efficiency of inspections, and provides support for the safe and stable operation of the power system.
Smart Images

Figure CN120234745A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular, to an intelligent electricity consumption behavior inspection method, device, and storage medium based on multi-source data. Background Art
[0002] Power supply enterprise electricity inspection is a key area for the power industry to ensure the order of power supply and use, and maintain economic benefits and social fairness. With the rapid growth of power demand and the complexity of electricity consumption behavior, the importance of electricity inspection has become increasingly prominent, which is directly related to the reasonable allocation of power resources and the operation efficiency of enterprises. However, traditional inspection methods often seem powerless when faced with massive data and diverse electricity consumption scenarios, which makes the research of more efficient and intelligent inspection means an urgent need for the industry development. Currently, many power supply enterprises rely on manual inspections or simple data statistics methods to identify abnormal electricity consumption behaviors. This method is not only inefficient but also easily interfered by human factors, resulting in frequent omissions or misjudgments. In addition, the existing methods are insufficient in integrating multi-source information and real-time analysis, and it is difficult to adapt to the dynamically changing electricity consumption environment. These limitations make it impossible for the inspection work to fully cover the growing electricity consumption data and respond to potential violations in a timely manner. In this context, the core challenges faced in the field of electricity inspection gradually emerge, and the most prominent technical factors include data real-time performance, abnormal detection accuracy, and multi-dimensional information fusion ability. First, the realization of real-time monitoring of electricity consumption data is limited by the processing speed and response ability of the existing system, and it is difficult to quickly respond to sudden electricity consumption anomalies. Second, the identification of abnormal electricity consumption behaviors relies on single-dimensional analysis, lacking comprehensive judgment of features such as sudden changes in electricity consumption and abnormal time periods, resulting in insufficient accuracy of the detection results. Finally, there are still technical barriers in the integration of multi-source data, such as the mining of the correlation between historical electricity consumption records, the status of metering devices, and user information, which directly affects the comprehensiveness of inspection and decision-making support capabilities. These unsolved technical factors together constitute a unique problem in improving inspection efficiency and accuracy. Therefore, how to achieve real-time monitoring, accurately identify abnormal behaviors, and effectively integrate multi-dimensional information in massive electricity consumption data has become a key issue in the digital transformation of power supply enterprise electricity inspection. The solution to this problem not only requires breaking through the existing technical bottlenecks but also constructing an inspection system that can support dynamic analysis and intelligent early warning, so as to promote the overall upgrade of the industry supervision ability. Summary of the Invention
[0003] To achieve the object of the present invention, in the first aspect, the present invention provides an intelligent electricity consumption behavior inspection method based on multi-source data, mainly including: The distributed computing framework is adopted to process the power consumption data stream collected in real time by sensors and smart meters. Through parallel computing technology, the integration and standardization of the power consumption data stream are completed to generate the original power consumption monitoring data set. For the dynamically updated original power consumption monitoring data set, the sliding window algorithm is applied to extract the power consumption data within a specified period, calculate the real-time change characteristics. If it is detected that the real-time change characteristics exceed the corresponding dynamic threshold, it is marked as an abnormal event to generate time series data. Among them, the time series data includes basic identification, time information, real-time change characteristics and abnormal marks. The basic identification includes device identification, user number, and meter model. From the time series data and historical power consumption records, multi-dimensional feature vectors are extracted. Through a pre-constructed classifier, the classification results of abnormal behaviors are output. The user registration information and the classification results of the abnormal behaviors are obtained to construct a user-abnormal event association matrix. Through the association rule mining algorithm, the user-abnormal event association matrix is analyzed to generate a multi-source association data set of abnormal behaviors. Based on the multi-source association data set of abnormal behaviors, real-time stream processing is performed on the newly incoming power consumption data. Through the incremental update mechanism, if the newly added data triggers the preset abnormal pattern matching condition, an abnormal behavior warning signal is immediately generated. The principal component analysis method is used to generate a high-risk event feature vector according to the abnormal behavior warning signal. The high-risk event feature vector, the original power consumption monitoring data set and the newly incoming power consumption data are integrated to obtain a fused data vector. The fused data vector is input into a convolutional neural network to generate the finally confirmed abnormal behavior data. Based on the finally confirmed abnormal behavior data, visualization technology is used to generate a real-time monitoring dashboard to dynamically display the abnormal behaviors of the power supply enterprise's electricity inspection.
[0004] In a second aspect, the present invention also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0005] In a third aspect, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.
[0006] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: The present invention discloses an intelligent power consumption behavior inspection method, device, and storage medium based on multi-source data. It processes a large amount of real-time power consumption data through distributed computing, applies a sliding window algorithm to extract power consumption features and mark abnormal events, constructs a multi-dimensional feature vector for abnormal classification in combination with historical records, mines abnormal behavior association rules by integrating user information, realizes real-time abnormal early warning for new data, and uses deep learning technology to confirm high-risk events. Finally, the abnormal detection results are displayed in a visual manner. This method can effectively integrate multi-source heterogeneous data, realize intelligent analysis and real-time monitoring of power consumption behavior, improve the accuracy and efficiency of power consumption inspection in power supply enterprises, and provide strong support for the safe and stable operation of the power system. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 It is a flowchart of the intelligent power consumption behavior inspection method based on multi-source data of the present invention.
[0008] Figure 2 It is a structural schematic block diagram of the computer device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0009] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0010] As Figure 1 , the intelligent power consumption behavior inspection method based on multi-source data in this embodiment may specifically include: Step S101: Use a distributed computing framework to process the power consumption data stream collected in real time by sensors and smart meters, and complete the integration and standardization of the power consumption data stream through parallel computing technology to generate an original power consumption monitoring data set.
[0011] Obtain the power consumption data stream collected in real time by sensors and smart meters through a distributed computing framework, use parallel computing technology to split the data stream to obtain a preliminarily processed data segment. Extract multi-source data from the preliminarily processed data segment, and use a distributed stream processing platform to unify the format of the multi-source data to obtain an integrated data set. For the integrated data set, if the field formats are inconsistent, use Pandas to adjust the field types and units to obtain the original power consumption monitoring data set.
[0012] Distributed computing frameworks are often used to process large-scale data streams, especially in scenarios of electricity consumption data monitoring. Exemplarily, a city power grid company has deployed a distributed computing framework based on Apache Spark to obtain real-time electricity consumption data streams collected by sensors and smart meters. The sensors collect data such as voltage and current per second, and the smart meters record the electricity consumption once per minute, with the total data volume reaching millions of records per hour. Through the Spark Streaming stream processing framework in Apache Spark, the system divides the data stream into time windows, for example, forming a batch every 5 seconds, and distributes it to cluster nodes for parallel processing to generate a preliminarily processed data segment. This method effectively reduces the pressure on a single node and improves real-time performance. In one possible implementation, when extracting multi-source data from the preliminarily processed data segment, the heterogeneity of sensor data and meter data needs to be processed. For example, sensor data contains high-frequency signals such as voltage and current, and the format is JSON; meter data is in CSV format and contains user IDs, electricity consumption, etc. Distributed stream processing platforms such as Apache Flink can be used to unify the format. Specifically, Apache Flink converts JSON and CSV data into Avro format by defining a unified event schema, generating an integrated data set. The Avro format supports dynamic expansion, reduces the overhead of format conversion, and ensures data consistency. It should be noted that if the field formats in the integrated data set are inconsistent, such as the voltage unit being volts or kilovolts, and the timestamp format being ISO8601 or Unix timestamp, further processing is required. In one embodiment, Pandas is used to adjust the field types and units in a distributed environment. For example, for a data subset processed by a certain node, Pandas uniformly converts kilovolts to volts and converts the timestamp to the standard UTC format, finally generating an original data set for electricity consumption monitoring. The fields in the original data set for electricity consumption monitoring are consistent, facilitating subsequent analysis and modeling. This standardization operation ensures the accuracy of subsequent analysis. Preferably, the effects brought by the above technical combination are significant. For example, the parallel computing of Apache Spark increases the data processing speed by about 3 times, the stream processing of Apache Flink ensures millisecond-level latency, and the field adjustment of Pandas reduces the data cleaning time by about 20%. It can be understood that these technologies work together to form a complete link from data collection to preprocessing, supporting real-time electricity consumption monitoring. For example, the power grid company can analyze peak electricity consumption periods based on the original data set, optimize power dispatching, and reduce the peak load pressure by about 15%. In one embodiment, considering the scenario of a sharp increase in data volume, the system can dynamically expand the Apache Spark cluster nodes from 10 to 20 to ensure that the processing capacity increases with the growth of the data volume. For example, during the peak electricity consumption period on holidays, the performance of processing millions of records per second can still be maintained after the node expansion.It should be noted that the fault tolerance mechanism of Apache Flink ensures the recoverability from the most recent checkpoint after the interruption of data stream processing and guarantees data integrity through the checkpoint technology. Specifically, when dealing with inconsistent fields, Pandas can also automatically identify outliers through a predefined rule library. For example, when the voltage value in a user's electricity meter data suddenly changes to 0, Pandas combines historical data to determine it as an anomaly and marks it, reducing the cost of manual verification. This not only improves data quality but also provides reliable input for subsequent machine learning model training. For example, based on the above dataset, the power grid company can further analyze the electricity consumption behavior of users, identify high-energy-consuming devices, and provide a basis for energy-saving strategies. It can be understood that the high efficiency of distributed computing, the real-time nature of stream processing, and the flexibility of Pandas jointly build a robust electricity consumption monitoring system, providing technical support for the refined management of smart grids.
[0013] Step S102: For the dynamically updated original dataset of electricity consumption monitoring, apply the sliding window algorithm to extract the electricity consumption data within a specified time period, calculate the real-time change characteristics. If it is detected that the real-time change characteristics exceed the corresponding dynamic threshold, mark it as an abnormal event and generate time series data. The time series data includes a basic identifier, time information, real-time change characteristics, and an abnormal mark. The basic identifier includes a device identifier, a user number, and a meter model.
[0014] Group the dynamically updated original dataset of electricity consumption monitoring by device identifier and timestamp, set the window size to 60 seconds and the step size to 10 seconds to divide it into multiple windows. Calculate the real-time change characteristics for the electricity consumption data within each window. The real-time change characteristics include the difference in electricity consumption at the beginning and end of the window, the change rate per unit time, the fluctuation amplitude between the maximum and minimum values within the window, and the year-on-year change rate compared with the previous window. Generate the dynamic thresholds corresponding to each real-time change characteristic according to the preset dynamic threshold rules. If any real-time change characteristic in the current window exceeds the corresponding dynamic threshold, mark it as an abnormal event and generate time series data including a basic identifier, time information, real-time change characteristics, and an abnormal mark.
[0015] In an electricity consumption monitoring system, dividing the original data set into time windows is the basis for anomaly detection. Exemplarily, a power company analyzes the data of smart electricity meters in a residential area. The device identifier is the meter ID, such as "SM-10086", and the timestamp is accurate to the second. The system uses 60 seconds as the window size and slides every 10 seconds to form overlapping windows. For example, the first window is from 08:00:00 to 08:01:00, the second window is from 08:00:10 to 08:01:10, and so on. In a possible implementation, the real-time change feature calculation process is specifically as follows: For the window from 08:00:00 to 08:01:00 of device "SM-10086", the electricity consumption at the beginning and end is 2.5 kWh and 2.8 kWh respectively, and the difference is 0.3 kWh; the change rate per unit time is 0.3 / 60 = 0.005 kWh / s; the fluctuation range between the maximum value of 2.9 kWh and the minimum value of 2.5 kWh within the window is 0.4 kWh; compared with the end value of 2.4 kWh in the previous window (from 07:59:50 to 08:00:50), the month-on-month change rate is (2.8 - 2.4) / 2.4 = 16.7%. It should be noted that the dynamic threshold rule is generated based on the statistical characteristics of historical data. Specifically, the system may calculate the mean and standard deviation of each feature according to the electricity consumption data in the same period of the past 30 days, and set the threshold range as the mean plus or minus 3 times the standard deviation. Data outside this range is marked as abnormal data. For example, the mean change rate per unit time in a certain period is 0.003 kWh / s, and the standard deviation is 0.0007 kWh / s, then the upper limit of the dynamic threshold is 0.003 + 3×0.0007 = 0.0051 kWh / s, and the lower limit is 0.003 - 3×0.0007 = 0.0009 kWh / s. This means that in this period, the change rate per unit time should normally be between 0.0009 - 0.0051 kWh / s. For the month-on-month change rate dynamic threshold, calculate the month-on-month change rate of the electricity consumption data in each time window, and calculate its mean and standard deviation. Determine the month-on-month change rate dynamic threshold by integrating the fluctuation characteristics of the electricity consumption data and the tolerance for anomalies in actual business. Exemplarily, the mean month-on-month change rate in a certain period is 10%, and the standard deviation is 2%. The dynamic threshold calculation formula is the mean plus 2.5 times the standard deviation, so the dynamic threshold is 10% + 2.5×2% = 15%. In an embodiment, if the change rate per unit time of 0.005 kWh / s of "SM-10086" in the window from 08:00:00 to 08:01:00 is close to but does not exceed the threshold of 0.0051 kWh / s, the system does not mark it as abnormal; however, if the month-on-month change rate of 16.7% exceeds the dynamic threshold of 15% for this feature, it is marked as an abnormal event. The generated time series data includes the basic identifier "SM-10086", time information "from 2023-05-15-08:00:00 to 08:01:00", real-time change feature values, and the "month-on-month change anomaly" mark.Preferably, this anomaly detection method based on multi-features and sliding windows can detect electricity anomalies in a timely manner, such as equipment failures or power theft behaviors, providing decision-making support for power grid management.
[0016] Step S103: Extract multi-dimensional feature vectors from the time series data and historical electricity consumption records, and output the classification results of abnormal behaviors through a pre-constructed classifier.
[0017] Extract time information, real-time change features, and basic identifiers from the time series data. Obtain the historical electricity consumption records of the device in the past 30 days through the basic identifiers, where the historical electricity consumption records include the average daily electricity consumption, the standard deviation of the load curve, and the average electricity consumption of similar devices in the same period. Concatenate the time information, real-time change features, and historical electricity consumption records into multi-dimensional features to generate multi-dimensional feature vectors. Input the multi-dimensional feature vectors into a classifier pre-constructed by the random forest algorithm, and output the classification results of abnormal behaviors. Among them, the classification results of preset abnormal behaviors include normal fluctuation results, suspected equipment failure results, abnormal user behavior results, and data acquisition abnormal results.
[0018] In an electricity consumption monitoring system, multi-dimensional feature analysis is the key to identifying abnormal electricity consumption behaviors. By extracting the time information, real-time change features, and basic identifiers from time-series data, the system can construct a complete electricity consumption profile. Exemplarily, the power monitoring system of an industrial park has collected data of a high-energy-consuming device labeled "DV-10086". The system first extracts the basic identifier information of this device, including the device identifier "DV-10086", the user number "U2023-567", and the electricity meter model "SM-500X". In a possible implementation, the system obtains the historical electricity consumption records of this device in the past 30 days and finds that its average daily electricity consumption is 1,250 kWh, the standard deviation of the load curve is 78.5, while the average electricity consumption of similar devices during the same period is 1,180 kWh. These historical data provide a benchmark reference for anomaly judgment. Specifically, when the system detects real-time change features (such as the change rate per unit time reaching 25 kW / minute, which is much higher than the normal value of 8 kW / minute) within a time window, it will splice these features with the historical electricity consumption records to form a feature vector containing multiple dimensions. It should be noted that the construction process of the multi-dimensional feature vector integrates time dimension information (such as weekday / non-weekday markers, time period types), real-time change features (such as fluctuation amplitude, change rate), and historical electricity consumption records (such as average daily electricity consumption, load curve features). Preferably, after receiving these feature vectors, a pre-trained random forest classifier can output accurate abnormal behavior classification results. For example, when the system detects that the device suddenly shows periodic severe fluctuations in electricity consumption during normal working hours and the fluctuation pattern is similar to the historical fault pattern, the classifier may output a "suspected device fault result" to remind the maintenance personnel to conduct an inspection. When the system detects a sudden increase in the electricity consumption of a certain device in the early morning, if it is marked as abnormal in step S102, but step S103 finds that there are load fluctuations caused by regular maintenance of this device during the same historical period, it is classified as a "normal fluctuation result". It can be understood that different types of anomalies have their unique characteristic manifestations. For example, the "abnormal user behavior result" usually shows a sudden large increase in electricity consumption without obvious fluctuations, while the "abnormal data collection result" may show data jumps or discontinuities.
[0019] Step S104, obtain the user registration information and the classification result of the abnormal behavior, construct a user-abnormal event association matrix, and analyze the user-abnormal event association matrix through an association rule mining algorithm to generate a multi-source association data set of abnormal behaviors.
[0020] Obtain the basic attributes in the user registration information to generate user basic portrait features, where the basic attributes include user type, geographical location, registration duration, and device type; read the data corresponding to the abnormal behavior classification results and extract the core features of abnormal behavior, where the core features include abnormal type, occurrence time, duration, and influence range; process the user basic portrait features and the core features of abnormal behavior through a data cleaning process, including removing missing values and outliers, and normalizing numerical features to obtain standardized user basic portrait features and core features of abnormal behavior; associate the standardized user basic portrait features and the core features of abnormal behavior according to the user ID to construct a user-abnormal event association matrix, where each row of the user-abnormal event association matrix represents a user and each column represents an abnormal feature; use the Apriori algorithm to analyze the user-abnormal event association matrix. If the confidence of the association rule is higher than the preset confidence threshold, record the rule in the association rule library. If the confidence of the association rule is lower than the preset confidence threshold, discard the rule; group the high-confidence association rules through hierarchical clustering to generate a multi-source association dataset of abnormal behavior.
[0021] In the business scenario of the power consumption monitoring system, generating association rules based on user registration information and abnormal behavior data and constructing a multi-source association dataset are important methods for optimizing anomaly detection. Exemplarily, obtaining the basic attributes in the user registration information is the first step in constructing a user portrait. The basic attributes include user type, geographical location, registration duration, and device type. The user type can be divided into residential users, commercial users, and industrial users; the geographical location refers to the area where the user is located, such as a city or an industrial park; the registration duration records the length of time the user has been accessing the system, such as 3 years; the device type refers to the category of power-consuming devices, such as high-energy-consuming devices or smart home appliances. In a possible implementation, the system extracts the information of an industrial user from the database: the user type is "industrial user", the geographical location is "a certain industrial park in a city", the registration duration is "5 years", and the device type is "heavy machinery". These attributes constitute the user's basic portrait features and provide an identity basis for subsequent association analysis. Specifically, reading the classification results of abnormal behavior and extracting the core features are the key to abnormal event analysis. The core features include abnormal type, occurrence time, duration, and impact scope. The abnormal type can be device failure, abnormal user behavior, or data collection anomaly; the occurrence time records the specific moment when the anomaly occurs, such as "14:00 on April 20, 2025"; the duration refers to the time the anomaly lasts, such as "2 hours"; the impact scope refers to the devices or areas affected, such as "a single device". For example, the system detects that a certain device has an "equipment failure" anomaly at the early morning of a working day, the occurrence time is "02:00 on April 21, 2025", and it lasts for 1 hour, only affecting that device. It should be noted that data cleaning ensures the quality of feature data. The cleaning process includes removing missing values and outliers, and normalizing numerical features. Missing values refer to incomplete data records, such as the geographical location of a certain user being empty; outliers refer to data that significantly deviates from the normal range, such as the registration duration being recorded as "999 years". Normalization maps numerical features such as registration duration and duration to the interval from 0 to 1. For example, the system discovers that the duration in a certain user record is "-1 hour", removes it after identifying it as an outlier, and normalizes the registration duration, mapping 5 years to 0.5. The cleaned feature data is standardized to improve the accuracy of subsequent analysis. In one embodiment, constructing a user-abnormal event association matrix requires associating the user's basic portrait features with the core features of abnormal behavior through the user ID. Each row of the user-abnormal event association matrix represents a user, and each column represents an abnormal feature, such as "equipment failure duration" or "abnormal occurrence time". For example, the system generates a row of data for user U2023-567, including its industrial user type, 5-year registration duration, and 2-hour duration of a certain equipment failure. The specific example is shown in the following table:
[0022] This matrix structure is clear, facilitating the exploration of the correlation between users and anomalies. Preferably, the Apriori algorithm is used to analyze the correlation matrix to mine high-confidence rules. The Apriori algorithm identifies frequent item sets and generates rules, such as "industrial users and registration duration > 3 years → equipment failure". The confidence threshold is set at 0.8. Rules with a value higher than this are recorded in the association rule library, and those lower are discarded. For example, when the system discovers that the rule "commercial users and geographical location is urban area → data collection anomaly" has a confidence of 0.85, it is retained. This method effectively filters out reliable anomaly patterns. It can be understood that hierarchical clustering groups high-confidence rules to generate a multi-source association data set. Hierarchical clustering clusters rules into several groups according to the similarity of rule features, such as "rules related to equipment failure" or "rules related to abnormal user behavior". For example, the system clusters "industrial users → equipment failure" and "registration duration > 5 years → equipment failure" into one group to generate a multi-source association data set for equipment failure. This grouping helps the system formulate precise response strategies for different anomaly types. Through the above methods, the system forms a complete chain from user portraits to anomaly features, and rule mining and grouping further enhance the pertinence of anomaly detection. The implementation of each step supports each other to ensure the reliability and practicality of the analysis results.
[0023] Step S105: Based on the multi-source association data set of abnormal behaviors, perform real-time stream processing on newly incoming power consumption data. Through the incremental update mechanism, if the new data triggers the preset anomaly pattern matching condition, an abnormal behavior warning signal is generated immediately.
[0024] Obtain the newly incoming power consumption data from the real-time data stream, extract the data stream timestamp and power consumption data features, and generate an initial data vector. Through the incremental update mechanism, compare the features of the initial data vector with the abnormal behavior classification results in the multi-source association data set of abnormal behaviors to determine the data similarity. If the data similarity is higher than the preset similarity threshold, determine the abnormal behavior classification result according to the preset abnormal event identifier to obtain an incremental abnormal behavior classification result. Use a data processing efficiency optimization method to generate an abnormal behavior warning signal based on the incremental abnormal behavior classification result.
[0025] Specifically, in the scenario of smart grid user behavior monitoring, the processing of real-time data streams is the core of abnormal behavior detection. Exemplarily, to obtain newly incoming power consumption data from real-time data streams, key fields in the data packets need to be parsed, such as timestamps and power consumption data features. The timestamp records the specific moment when the data is generated, for example, 2025-04-26-22:30:00, which is used to track the data time sequence. The power consumption data features include instantaneous power, power consumption, and voltage fluctuations, etc. For example, the data stream of a certain user shows that the instantaneous power is 2.5 kilowatts and the voltage fluctuation amplitude is 5 volts. These features are extracted through a streaming processing framework to form an initial data vector with a structure of [timestamp, instantaneous power, voltage fluctuation], such as [2025-04-26-22:30:00, 2.5, 5]. In a possible implementation, an incremental update mechanism is used to compare the features of the initial data vector with the classification results of abnormal behaviors in a multi-source associated dataset. It should be noted that the multi-source associated dataset contains historical abnormal patterns, such as sudden power increase or voltage abnormality, etc., and each pattern corresponds to a feature range. During the comparison, the system calculates the similarity between the initial data vector and the abnormal pattern, for example, measured by the Euclidean distance. Assuming that the similarity threshold is 0.8 and the similarity of a certain data vector is 0.85, which is higher than the similarity threshold, it indicates that there may be an abnormality. Specifically, the system determines the classification result as an abnormal behavior according to the abnormal event identifier, such as "sudden power increase", and generates an incremental classification result, such as "User ID_002, sudden power increase, 2025-04-26-22:30:00". Preferably, the data processing efficiency optimization method improves the performance through streaming computing and caching mechanisms. For example, the system caches frequently accessed abnormal patterns in memory to reduce the database query time. In one embodiment, an abnormal behavior warning signal is generated based on the incremental classification result, and the signal contains the user ID, abnormal type, and timestamp. For example, for User ID_002, a warning signal "sudden power increase, 2025-04-26-22:30:00" is generated and sent to the monitoring platform through a message queue. It can be understood that this method supports real-time response and facilitates quick intervention. In another embodiment, the feature comparison can be supplemented from multiple aspects. For example, in addition to power and voltage, the power consumption period feature is added. For example, late-night power consumption may be indirect evidence of an abnormality. Assuming that the power consumption data of User ID_002 suddenly increases at night, combined with historical data analysis, the system further confirms the possibility of an abnormality. This multi-dimensional comparison improves the classification accuracy. For example, the historical data of a certain user shows that their nighttime power consumption is usually less than 0.5 kilowatts, while the current data is 2.5 kilowatts, triggering a more confident abnormal determination. It should be noted that after the warning signal is generated, it can be graded according to the severity of the abnormality. For example, a sudden power increase lasting for 5 minutes is a low-level abnormality, and lasting for 1 hour is a high-level abnormality. The system adjusts the warning priority accordingly, such as immediately pushing high-level abnormalities to the administrator.Exemplarily, for a continuous 1-hour power surge for user ID_002, a high-level warning signal is generated, accompanied by detailed abnormal characteristics for subsequent analysis and processing. This grading mechanism optimizes resource allocation and improves monitoring efficiency.
[0026] Step S106: Use the principal component analysis method to generate a high-risk event feature vector based on the abnormal behavior warning signal. Integrate the high-risk event feature vector, the original power consumption monitoring data set, and the newly incoming power consumption data to obtain a fused data vector. Input the fused data vector into a convolutional neural network to generate the finally confirmed abnormal behavior data.
[0027] Obtain the abnormal behavior warning signal, and use the principal component analysis method to generate a high-risk event feature vector based on the abnormal behavior warning signal. Obtain the original power consumption monitoring data set and the newly incoming power consumption data, and use a data cleaning tool to process the data to generate a multi-source data set. Use the weighted average method to integrate the high-risk event feature vector and the multi-source data set to obtain a fused data vector. Implement a convolutional neural network through the TensorFlow framework, input the fused data vector into the convolutional neural network to generate the finally confirmed abnormal behavior data, where the finally confirmed abnormal behavior data includes abnormal power consumption records and abnormal types.
[0028] In the scenario of monitoring user behavior in the smart grid, the processing and analysis of abnormal behavior warning signals are important links to ensure power consumption safety. Exemplarily, after obtaining the abnormal behavior warning signal, the principal component analysis method can be used to generate a high-risk event feature vector. The principal component analysis method is a dimensionality reduction technique aimed at extracting the most representative features in the data. The abnormal behavior warning signal includes data such as user information, real-time change features, abnormal types, and association rule information. Suppose a warning signal contains a user ID, an abnormal type such as "sudden power increase", and a timestamp, such as "user ID_003, sudden power increase, 2025-04-27-01:00:00". Through principal component analysis, the system extracts key features from the signal, such as the abnormal duration and the power change amplitude, and generates a feature vector, such as [duration: 30 minutes, power change: 3 kW], which is convenient for subsequent analysis. In a possible implementation, after obtaining the original power consumption monitoring data set and the newly incoming power consumption data, a data cleaning tool needs to be used for processing. The original data set may contain missing values or noisy data, such as the voltage data of a certain user being empty. The data cleaning tool can fill in the missing values through interpolation or filter out the outliers, such as removing the data with voltage values outside the reasonable range. After cleaning, a multi-source data set is generated, which includes features such as timestamps, instantaneous power, and power consumption periods. For example, a multi-source data set records that the instantaneous power of user ID_003 at 2025-04-27-01:00:00 is 3.2 kW and the power consumption period is late at night, forming a multi-source data set. Specifically, the weighted average method is used to integrate the high-risk event feature vector and the multi-source data set to generate a fused data vector. The weighted average method assigns weights according to the feature importance and normalizes the high-risk event feature vector and the multi-source data set. The specific example is as follows: High-risk event feature vector: Suppose it contains "abnormal duration" and "power change amplitude", expressed as V1 = [duration = 30 minutes, power change = 3 kW], and after normalization, it is V1norm = [0.6, 0.4]; Multi-source data set features: It contains "instantaneous power" and "power consumption period", expressed as V2 = [instantaneous power = 3.2 kW, power consumption period (late at night) = 1], and after normalization, it is V2norm = [0.5, 0.5]; According to the preset weights, the high-risk feature weight is α = [0.7, 0.3], and the multi-source data weight is β = [0.3, 0.7]; Fused data vector = α×V1norm + β×V2norm = [0.7×0.6 + 0.3×0.5, 0.7×0.4 + 0.3×0.5] = [0.57, 0.43]. This method ensures the effective combination of multi-dimensional information. Preferably, a convolutional neural network is implemented through the TensorFlow framework, and the fused data vector is input into the network to generate the finally confirmed abnormal behavior data. The convolutional neural network is good at capturing local patterns in the data. In one embodiment, the network inputs the fused data vector and outputs abnormal power consumption records and abnormal types.For example, for user ID_003, after network analysis, the abnormal record is confirmed as "2025-04-27-01:00:00, power suddenly increased by 3.2 kW", and the abnormal type is "abnormal electricity consumption at night". It can be understood that the network extracts deep features through multi-layer convolution and pooling to improve the classification accuracy. It should be noted that multi-faceted analysis can further support the results. For example, combined with the user's historical electricity consumption habits, the system finds that the user ID_003's electricity consumption at night is usually less than 0.5 kW, and the current 3.2 kW is significantly higher, increasing the abnormal confidence level. In addition, the analysis of the abnormal duration shows that the 30-minute sudden increase conforms to the intermediate abnormal mode. These aspects jointly confirm the reliability of the abnormal behavior. Exemplarily, if the abnormal type is "abnormal electricity consumption at night", the system can generate a detailed record, including the specific time and features, for subsequent tracking. This way of combining multi-dimensional analysis and deep learning can effectively identify abnormal behaviors in complex scenarios.
[0029] Step S107, based on the finally confirmed abnormal behavior data, use visualization technology to generate a real-time monitoring dashboard to dynamically display the abnormal behaviors of the power supply enterprise's electricity inspection.
[0030] Extract the abnormal electricity consumption records confirmed within the last 30 days from the finally confirmed abnormal behavior data, classify and mark them according to user type, geographical location, and abnormal type to obtain a structured abnormal data set. Calculate the key performance indicators according to the structured abnormal data set, and use the sliding window algorithm to update the indicator values in real time. Among them, the key performance indicators include the abnormal detection accuracy index, the early warning coverage percentage, and the average response time. Build a multi-level interactive dashboard through the D3.js visualization library, set the indicator threshold warning line, and if the indicator value exceeds the warning line, trigger a visual warning signal to display the heat map of the abnormal area. Establish a time series analysis model for the dashboard data, use the exponential smoothing method to predict the change trend of each indicator within the next 24 hours, and generate a trend prediction curve. Combine the geographical information system to draw a distribution map of abnormal events, automatically cluster according to the abnormal density to form high-risk area marks, and display the inspection resource allocation suggestions for each area. Use the WebSocket technology to achieve real-time data push. When a new abnormal event is confirmed, automatically update the indicators and charts of the dashboard to keep the monitoring data real-time.
[0031] In the scenario of monitoring the behavior of smart grid users, extracting the abnormal power consumption records confirmed within the most recent 30 days from the finally confirmed abnormal behavior data and classifying and labeling them are the key steps in optimizing the monitoring system. Exemplarily, the system filters out the records with a time range from March 28, 2025 to April 26, 2025 from the abnormal behavior data, including information such as user ID, abnormal type, timestamp, and geographical location. The user types can be divided into residential users, commercial users, and industrial users; the geographical location is divided by administrative region, such as "Chaoyang District, Beijing"; the abnormal types include "sudden increase in power", "voltage anomaly", etc. In one possible implementation, the system extracts the records through database queries and uses a classification algorithm to generate a structured abnormal data set according to the above dimensions. For example, residential user ID_005 had a sudden increase in power at 23:00:00 on April 25, 2025 in Chaoyang District, and was labeled as "residential - Chaoyang District - sudden increase in power". Specifically, key performance indicators are calculated based on the structured abnormal data set, and the sliding window algorithm is used to update the indicator values in real time. The key performance indicators include the abnormal detection accuracy index, the warning coverage percentage, and the average response time. The sliding window algorithm uses a 24-hour window and slides once per hour to count the abnormal events within the window. Preferably, the abnormal detection accuracy index is calculated by comparing the confirmed abnormal events with the actual inspection results. For example, within a certain window, 100 abnormal events are confirmed, and 90 are correctly verified by inspection, with an accuracy of 90%. The warning coverage percentage measures the proportion of abnormal events covered by the warning signals. For example, on a certain day, 80 warnings are issued, covering 70 real abnormal events, with a coverage rate of 87.5%. The average response time records the duration from the warning being issued to the inspection response. For example, it is 15 minutes on average. In one embodiment, a multi-level interactive dashboard is constructed through the D3.js visualization library to display the above indicators and set threshold warning lines. The dashboard includes bar charts, line charts, and heat maps, intuitively presenting the changes in the indicators. For example, the accuracy threshold is set to 85%. If the accuracy drops to 80% in a certain hour, the dashboard triggers a red visual warning signal, and at the same time, the heat map highlights the high-abnormal regions such as Chaoyang District. It should be noted that the heat map is dynamically colored according to the density of abnormal events, and the regions with high density are shown in dark red, facilitating the quick positioning of problem areas. It can be understood that a time series analysis model is established for the dashboard data, and the exponential smoothing method is used to predict the change trend of the indicators within the next 24 hours. The exponential smoothing method predicts the future trend by weighting the historical indicator values. For example, the system analyzes the accuracy data of the past 7 days and predicts that the accuracy may drop from 90% to 88% within the next 24 hours, generating a trend prediction curve to assist managers in adjusting strategies in advance. In one embodiment, an abnormal event distribution map is drawn in combination with the geographic information system. The system generates a distribution map based on the longitude and latitude data of the abnormal events and identifies high-risk areas through the density clustering algorithm. For example, a certain community in Chaoyang District has the highest density of abnormal events and is marked as a high-risk area, and the system recommends increasing the number of inspection personnel to 5 people.For example, the WebSocket technology is used to achieve real-time data push, ensuring that the dashboard is updated immediately after a new exception event is confirmed. Suppose a new exception record of a commercial user is added at 01:00:00 on April 26, 2025. The WebSocket pushes the data to the front end, and the exception count and heat map on the dashboard are refreshed in real time, and the accuracy index is updated to 89%. This real-time nature ensures the rapid response of the monitoring system to exception events and improves the efficiency of power grid safety management.
[0032] Referring to Figure 2 , an embodiment of the present application also provides a computer device, which can be a server, and its internal structure can be as Figure 2 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. The processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as the original dataset of power consumption monitoring and the classification results of abnormal behaviors. The network interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements the personalized recommendation method based on real-time user behavior in any of the above embodiments.
[0033] Those skilled in the art can understand that Figure 2 the structure shown in
[0034] is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied.
[0035] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0036] As described above, this is only the specific implementation manner of this specification. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here. It should be understood that the protection scope of this specification is not limited thereto. Any person skilled in the art within the technical scope disclosed in this specification can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of this specification.
Claims
1. An intelligent electricity consumption behavior inspection method based on multi-source data, characterized in that The method includes: Using a distributed computing framework to process the power consumption data stream collected in real time by sensors and smart meters, integrating and standardizing the power consumption data stream through parallel computing technology to generate an original power consumption monitoring data set; for the dynamically updated original power consumption monitoring data set, applying a sliding window algorithm to extract the power consumption data within a specified period, calculating the real-time change characteristics, if it is detected that the real-time change characteristics exceed the corresponding dynamic threshold, marking it as an abnormal event, generating time series data, where the time series data includes a basic identifier, time information, real-time change characteristics and an abnormal mark, and the basic identifier includes a device identifier, a user number, and a meter model; extracting multi-dimensional feature vectors from the time series data and historical power consumption records, and outputting the classification result of the abnormal behavior through a pre-constructed classifier; obtaining user registration information and the classification result of the abnormal behavior, constructing a user-abnormal event association matrix, analyzing the user-abnormal event association matrix through an association rule mining algorithm, and generating a multi-source association data set of abnormal behavior; based on the multi-source association data set of abnormal behavior, performing real-time stream processing on the newly incoming power consumption data, through an incremental update mechanism, if the newly added data triggers a preset abnormal pattern matching condition, immediately generating an abnormal behavior warning signal; using the principal component analysis method to generate a high-risk event feature vector according to the abnormal behavior warning signal, integrating the high-risk event feature vector, the original power consumption monitoring data set and the newly incoming power consumption data to obtain a fused data vector, and inputting the fused data vector into a convolutional neural network to generate the finally confirmed abnormal behavior data; based on the finally confirmed abnormal behavior data, using visualization technology to generate a real-time monitoring dashboard to dynamically display the abnormal behavior of the power consumption inspection of the power supply enterprise.
2. The method according to claim 1, wherein The using of a distributed computing framework to process the power consumption data stream collected in real time by sensors and smart meters, integrating and standardizing the power consumption data stream through parallel computing technology to generate an original power consumption monitoring data set includes: Obtaining the power consumption data stream collected in real time by sensors and smart meters through a distributed computing framework, and using parallel computing technology to split the data stream to obtain a preliminarily processed data segment; Extracting multi-source data from the preliminarily processed data segment, and using a distributed stream processing platform to unify the format of the multi-source data to obtain an integrated data set; For the integrated data set, if the field formats are inconsistent, using Pandas to adjust the field types and units to obtain the original power consumption monitoring data set.
3. The method according to claim 1, characterized in that, The applying of a sliding window algorithm to the dynamically updated original power consumption monitoring data set to extract the power consumption data within a specified period, calculate the real-time change characteristics, if it is detected that the real-time change characteristics exceed the corresponding dynamic threshold, mark it as an abnormal event, and generate time series data, where the time series data includes a basic identifier, time information, real-time change characteristics and an abnormal mark, and the basic identifier includes a device identifier, a user number, and a meter model, includes: Grouping the dynamically updated original power consumption monitoring data set by device identifier and timestamp, setting the window size to 60 seconds and the step size to 10 seconds, and dividing it into multiple windows; For the electricity consumption data within each window, calculate the real-time change features, where the real-time change features include the difference in electricity consumption at the beginning and end of the window, the change rate per unit time, the fluctuation range between the maximum and minimum values within the window, and the month-on-month change rate compared with the previous window; Generate dynamic thresholds corresponding to each real-time change feature according to the preset dynamic threshold rules. If any real-time change feature in the current window exceeds the corresponding dynamic threshold, it is marked as an abnormal event, and time-series data including the basic identifier, time information, real-time change features, and abnormal marks is generated.
4. The method according to claim 1, wherein Extract multi-dimensional feature vectors from the time-series data and historical electricity consumption records, and output the classification results of abnormal behaviors through a pre-constructed classifier, including: Extract time information, real-time change features, and basic identifiers from the time-series data; Obtain the historical electricity consumption records of the device in the past 30 days through the basic identifier, where the historical electricity consumption records include the average daily electricity consumption, the standard deviation of the load curve, and the average electricity consumption of similar devices during the same period; Concatenate the time information, real-time change features, and historical electricity consumption records into multi-dimensional features to generate multi-dimensional feature vectors; Input the multi-dimensional feature vectors into a classifier pre-constructed by the random forest algorithm to output the classification results of abnormal behaviors. Among them, the preset classification results of abnormal behaviors include normal fluctuation results, suspected device failure results, abnormal user behavior results, and data acquisition abnormal results.
5. The method according to claim 1, characterized in that, Obtain the user registration information and the classification results of the abnormal behaviors, construct a user-abnormal event association matrix, and analyze the user-abnormal event association matrix through an association rule mining algorithm to generate a multi-source association dataset of abnormal behaviors, including: Obtain the basic attributes in the user registration information to generate user basic portrait features, where the basic attributes include user type, geographical location, registration duration, and device type; Read the data corresponding to the classification results of abnormal behaviors and extract the core features of abnormal behaviors, where the core features include abnormal type, occurrence time, duration, and influence range; Process the user basic portrait features and the core features of abnormal behaviors through a data cleaning process, including removing missing values and outliers, and normalizing numerical features to obtain standardized user basic portrait features and core features of abnormal behaviors; Associate the standardized user basic portrait features and the core features of abnormal behaviors according to the user ID to construct a user-abnormal event association matrix, where each row of the user-abnormal event association matrix represents a user and each column represents an abnormal feature; Use the Apriori algorithm to analyze the user-abnormal event association matrix. If the confidence level of the association rule is higher than the preset confidence threshold, record the rule in the association rule library. If the confidence level of the association rule is lower than the preset confidence threshold, discard the rule; Group the high-confidence association rules through hierarchical clustering methods to generate a multi-source association dataset of abnormal behaviors.
6. The method according to claim 1, characterized in that, Based on the multi-source association dataset of abnormal behaviors, perform real-time stream processing on newly incoming electricity consumption data. Through the incremental update mechanism, if the new data triggers the preset abnormal pattern matching condition, an abnormal behavior warning signal is generated immediately, including: Obtain newly incoming power consumption data from the real-time data stream, extract the data stream timestamp and power consumption data features, and generate an initial data vector; Through the incremental update mechanism, compare the features of the initial data vector with the abnormal behavior classification results in the multi-source associated dataset of abnormal behaviors to determine the data similarity; If the data similarity is higher than the preset similarity threshold, determine the abnormal behavior classification result according to the preset abnormal event identifier to obtain the incremental abnormal behavior classification result; Adopt a data processing efficiency optimization method to generate an abnormal behavior warning signal based on the incremental abnormal behavior classification result.
7. The method according to claim 1, characterized in that, Using the principal component analysis method to generate a high-risk event feature vector according to the abnormal behavior warning signal, integrate the high-risk event feature vector, the original power consumption monitoring dataset, and the newly incoming power consumption data to obtain a fused data vector, and input the fused data vector into a convolutional neural network to generate the finally confirmed abnormal behavior data, including: Obtain the abnormal behavior warning signal, and use the principal component analysis method to generate a high-risk event feature vector according to the abnormal behavior warning signal; Obtain the original power consumption monitoring dataset and the newly incoming power consumption data, and use a data cleaning tool to process the data to generate a multi-source data set; Adopt the weighted average method to integrate the high-risk event feature vector and the multi-source data set to obtain a fused data vector; Implement a convolutional neural network through the TensorFlow framework, input the fused data vector into the convolutional neural network to generate the finally confirmed abnormal behavior data, where the finally confirmed abnormal behavior data includes abnormal power consumption records and abnormal types.
8. The method according to claim 1, wherein Based on the finally confirmed abnormal behavior data, use visualization technology to generate a real-time monitoring dashboard to dynamically display the abnormal behaviors of the power supply enterprise's electricity inspection, including: Extract the abnormal power consumption records confirmed within the last 30 days from the finally confirmed abnormal behavior data, classify and mark them according to user type, geographical location, and abnormal type to obtain a structured abnormal data set; Calculate the key performance indicators according to the structured abnormal data set, and use the sliding window algorithm to update the indicator values in real time, where the key performance indicators include the abnormal detection accuracy index, the warning coverage percentage, and the average response time; Build a multi-level interactive dashboard through the D3.js visualization library, set the indicator threshold warning line, and if the indicator value exceeds the warning line, trigger a visual warning signal to display the heat map of the abnormal area; Establish a time series analysis model for the dashboard data, and use the exponential smoothing method to predict the change trend of each indicator within the next 24 hours to generate a trend prediction curve; Combine the geographic information system to draw a distribution map of abnormal events, automatically cluster according to the abnormal density to form high-risk area marks, and display the inspection resource allocation suggestions for each area; Use the WebSocket technology to realize real-time data push. When a new abnormal event is confirmed, automatically update the indicators and charts of the dashboard to keep the monitoring data real-time.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Group behavior power distribution network power consumption abnormity detection method based on user attribute tags
CN110288383A
Anti-electricity-stealing inspection monitoring method and anti-electricity-stealing inspection monitoring system
CN113239087A
Electric power marketing digital inspection method and system
CN119513760A
Movable Wall Printing System
KR102405574B1
Cited By
Electric energy meter data potential safety hazard analysis method and system
CN120805126A
Electricity consumption abnormity analysis method and device, electronic equipment and readable storage medium
CN120908585A
Power system electricity selling deviation monitoring control method
CN120912375A
A power system power selling deviation monitoring control method
CN120912375B
Guide rail curvature online detection method and system
CN121475107A