Intelligent electricity consumption behavior inspection methods, equipment, and storage media based on multi-source data

By processing electricity consumption data streams through a distributed computing framework and deep learning technology, constructing multi-dimensional feature vectors for anomaly detection, generating real-time early warning signals and displaying abnormal behavior, the problem of low efficiency in existing technologies is solved, and the efficiency and accuracy of intelligent electricity consumption behavior inspection are achieved.

CN120234745BActive Publication Date: 2025-12-02GUANGDONG TOPWAY NETWORK +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510704843.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-12-02
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Traditional electricity inspection methods are inefficient, susceptible to human interference, difficult to monitor massive amounts of data in real time, and lack the ability to integrate multi-dimensional information, resulting in insufficient accuracy in detecting abnormal electricity use and an inability to respond promptly to potential violations.

Method used

A distributed computing framework is used to process electricity data streams. A sliding window algorithm is applied to extract real-time changing features. A multi-dimensional feature vector is constructed by combining historical records. Abnormal behavior analysis is performed through association rule mining and deep learning techniques to generate real-time early warning signals. The abnormal detection results are displayed using visualization techniques.

Benefits of technology

It enables real-time monitoring and accurate identification of abnormal behavior from massive amounts of electricity consumption data, improves the efficiency and accuracy of electricity inspection, supports dynamic analysis and intelligent early warning, and enhances the regulatory capabilities of power supply companies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234745B_ABST
    Figure CN120234745B_ABST
Patent Text Reader

Abstract

This invention relates to the field of information technology, providing a method, device, and storage medium for intelligent electricity consumption behavior inspection based on multi-source data. The method includes: processing massive electricity consumption data streams collected in real time from sensors and smart meters using a distributed computing framework; integrating and standardizing the multi-source data to generate a dynamically updated raw dataset for electricity consumption monitoring; extracting multi-dimensional feature vectors from time-series data and historical electricity consumption records containing basic identifiers, time information, real-time change characteristics, and anomaly markers; outputting classification results of abnormal behaviors; extracting the finally confirmed abnormal behavior data; and using visualization technology to generate a real-time monitoring dashboard to dynamically display the abnormal behaviors detected by the power supply company's electricity consumption inspection. This invention effectively integrates multi-source heterogeneous data, enabling intelligent analysis and real-time monitoring of electricity consumption behavior, improving the accuracy and efficiency of power supply company's electricity consumption inspection, and providing strong support for the safe and stable operation of the power system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology, and in particular to a method, device, and storage medium for intelligent electricity consumption behavior inspection based on multi-source data. Background Technology

[0002] Electricity consumption inspection by power supply companies is a crucial area for ensuring power supply and consumption order, maintaining economic efficiency, and upholding social equity in the power industry. With the rapid growth of electricity demand and the increasing complexity of electricity consumption behavior, the importance of electricity consumption inspection is becoming increasingly prominent, directly impacting the rational allocation of power resources and the operational efficiency of enterprises. However, traditional inspection methods often fall short when faced with massive amounts of data and diverse electricity consumption scenarios, making the research into more efficient and intelligent inspection methods an urgent need for industry development. Currently, many power supply companies rely on manual inspections or simple data statistics methods to identify abnormal electricity consumption behavior. This approach is not only inefficient but also susceptible to human error, leading to frequent omissions or misjudgments. Furthermore, existing methods lack the ability to integrate multi-source information and perform real-time analysis, making it difficult to adapt to dynamically changing electricity consumption environments. These limitations prevent inspection work from comprehensively covering the ever-increasing electricity consumption data and from responding promptly to potential violations. Against this backdrop, the core challenges facing the field of electricity consumption inspection are gradually emerging, with the most prominent technical factors including data real-time performance, anomaly detection accuracy, and the ability to integrate multi-dimensional information. First, the real-time monitoring of electricity consumption data is limited by the processing speed and response capabilities of existing systems, making it difficult to react quickly to sudden electricity consumption anomalies. Second, the identification of abnormal electricity consumption behavior relies on single-dimensional analysis, lacking comprehensive judgment on characteristics such as sudden changes in electricity consumption and abnormal time periods, resulting in insufficient accuracy of detection results. Finally, the integration of multi-source data, such as the correlation mining between historical electricity consumption records, metering device status, and user information, still faces technical barriers, directly affecting the comprehensiveness of audits and decision support capabilities. These unresolved technical factors collectively constitute a unique challenge to improving audit efficiency and accuracy. Therefore, how to achieve real-time monitoring, accurate identification of abnormal behavior, and effective integration of multi-dimensional information from massive amounts of electricity consumption data has become a key issue in the digital transformation of electricity consumption audits for power supply companies. Solving this problem requires not only overcoming existing technical bottlenecks but also building an audit system capable of supporting dynamic analysis and intelligent early warning, thereby promoting a comprehensive upgrade of industry regulatory capabilities. Summary of the Invention

[0003] To achieve the objectives of this invention, in a first aspect, this invention provides an intelligent electricity consumption behavior inspection method based on multi-source data, mainly comprising:

[0004] A distributed computing framework is employed to process real-time electricity consumption data streams collected by sensors and smart meters. Parallel computing technology is used to integrate and standardize the electricity consumption data streams, generating a raw electricity consumption monitoring dataset. For the dynamically updated raw electricity consumption monitoring dataset, a sliding window algorithm is applied to extract electricity consumption data within a specified time period, calculate real-time change characteristics, and if the detected real-time change characteristics exceed the corresponding dynamic threshold, they are marked as abnormal events, generating time-series data. The time-series data includes basic identifiers, time information, real-time change characteristics, and abnormal markers. The basic identifiers include device identifier, user number, and meter model. Multi-dimensional feature vectors are extracted from the time-series data and historical electricity consumption records, and a pre-built classifier is used to output the classification results of abnormal behavior. User registration information and the classification results of the abnormal behavior are obtained to construct... The user-abnormal event association matrix is ​​analyzed using association rule mining algorithms to generate a multi-source association dataset of abnormal behaviors. Based on this dataset, newly input electricity consumption data is processed in real-time. Through an incremental update mechanism, if new data triggers preset abnormal pattern matching conditions, an abnormal behavior warning signal is generated immediately. Principal component analysis is used to generate high-risk event feature vectors based on the warning signals. These vectors, along with the original electricity consumption monitoring dataset and the newly input electricity consumption data, are integrated to obtain a fused data vector. This fused data vector is then input into a convolutional neural network to generate the final confirmed abnormal behavior data. Based on this data, a real-time monitoring dashboard is generated using visualization technology to dynamically display the abnormal behaviors identified during the power supply company's electricity consumption audits.

[0005] In a second aspect, the present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the methods described above.

[0006] Thirdly, a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described above.

[0007] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:

[0008] This invention discloses an intelligent electricity consumption behavior inspection method, device, and storage medium based on multi-source data. It processes massive amounts of real-time electricity consumption data through distributed computing, applies a sliding window algorithm to extract electricity consumption features and mark abnormal events, constructs multi-dimensional feature vectors based on historical records for anomaly classification, integrates user information to mine association rules for abnormal behaviors, achieves real-time anomaly warnings for new data, and utilizes deep learning technology to confirm high-risk events. Finally, the anomaly detection results are displayed in a visual manner. This method effectively integrates multi-source heterogeneous data, enabling intelligent analysis and real-time monitoring of electricity consumption behavior, improving the accuracy and efficiency of electricity consumption inspections for power supply companies, and providing strong support for the safe and stable operation of the power system. Attached Figure Description

[0009] Figure 1 This is a flowchart of the intelligent electricity consumption behavior inspection method based on multi-source data according to the present invention.

[0010] Figure 2 This is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention. Detailed Implementation

[0011] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0012] like Figure 1 The intelligent electricity consumption behavior auditing method based on multi-source data in this embodiment may specifically include:

[0013] Step S101: A distributed computing framework is used to process the real-time electricity consumption data streams collected by sensors and smart meters. Parallel computing technology is used to integrate and standardize the electricity consumption data streams to generate the original dataset for electricity consumption monitoring.

[0014] Real-time electricity consumption data streams collected by sensors and smart meters are acquired using a distributed computing framework. Parallel computing techniques are employed to segment the data streams, resulting in pre-processed data segments. Multi-source data is extracted from these segments, and a distributed stream processing platform is used to unify the format of the multi-source data, yielding an integrated dataset. For the integrated dataset, if field formats are inconsistent, Pandas is used to adjust field types and units, resulting in the original electricity consumption monitoring dataset.

[0015] Distributed computing frameworks are commonly used to process large-scale data streams, particularly excelling in electricity data monitoring scenarios. For example, a city power grid company deployed a distributed computing framework based on Apache Spark to acquire real-time electricity data streams collected by sensors and smart meters. Sensors collect voltage, current, and other data every second, while smart meters record electricity consumption once per minute, resulting in millions of records per hour. Using the Spark Streaming framework in Apache Spark, the system segments the data stream by time windows, for example, forming batches every 5 seconds, and distributes them to cluster nodes for parallel processing, generating initially processed data segments. This approach effectively reduces the pressure on single nodes and improves real-time performance. In one possible implementation, when extracting multi-source data from the initially processed data segments, the heterogeneity of sensor and meter data needs to be addressed. For example, sensor data contains high-frequency signals such as voltage and current, in JSON format; meter data is in CSV format, containing user IDs, electricity consumption, etc. Distributed stream processing platforms such as Apache Flink can be used to unify the format. Specifically, Apache Flink defines a unified event pattern to convert JSON and CSV data into Avro format, generating an integrated data set. The Avro format supports dynamic expansion, reducing the overhead of format conversion and ensuring data consistency. It should be noted that if the field formats of the integrated dataset are inconsistent, such as voltage units being volts or kilovolts, or timestamps being ISO8601 or Unix timestamps, further processing is required. In one embodiment, Pandas is used to adjust field types and units in a distributed environment. For example, for a subset of data processed by a node, Pandas uniformly converts kilovolts to volts and timestamps to standard UTC format, ultimately generating the original electricity monitoring dataset. Consistent fields in the original electricity monitoring dataset facilitate subsequent analysis and modeling. This standardization ensures the accuracy of subsequent analysis. Preferably, the combined effect of the above technologies is significant. For example, Apache Spark's parallel computing increases data processing speed by approximately 3 times, Apache Flink's stream processing guarantees millisecond-level latency, and Pandas' field adjustments reduce data cleaning time by approximately 20%. Understandably, these technologies work together to form a complete chain from data acquisition to preprocessing, supporting real-time electricity monitoring. For example, power grid companies can analyze peak electricity demand periods based on raw datasets to optimize power dispatch and reduce peak load pressure by approximately 15%. In one embodiment, considering scenarios with a surge in data volume, the system can dynamically expand the Apache Spark cluster nodes from 10 to 20, ensuring that processing capacity increases with data volume. For instance, during peak electricity demand periods such as holidays, the expanded nodes can still maintain a performance of processing millions of records per second.It's worth noting that Apache Flink's fault tolerance mechanism, through checkpointing, ensures that data stream processing can resume from the most recent checkpoint after an interruption, guaranteeing data integrity. Specifically, Pandas can automatically identify outliers using a predefined rule base when handling inconsistencies in fields. For example, if a user's electricity meter data suddenly drops to 0, Pandas, combined with historical data, identifies this as an anomaly and marks it, reducing manual verification costs. This not only improves data quality but also provides reliable input for subsequent machine learning model training. For instance, based on the aforementioned dataset, power grid companies can further analyze user electricity consumption behavior, identify high-energy-consuming equipment, and provide a basis for energy-saving strategies. Understandably, the efficiency of distributed computing, the real-time nature of stream processing, and the flexibility of Pandas collectively construct a robust electricity monitoring system, providing technical support for the refined management of smart grids.

[0016] Step S102: For the dynamically updated raw dataset of electricity consumption monitoring, apply the sliding window algorithm to extract the electricity consumption data within a specified time period, calculate the real-time change characteristics, and if the real-time change characteristics exceed the corresponding dynamic threshold, mark it as an abnormal event and generate time-series data. The time-series data includes basic identifiers, time information, real-time change characteristics, and abnormal markers. The basic identifiers include device identifiers, user numbers, and meter models.

[0017] The dynamically updated raw dataset of electricity consumption monitoring is grouped by device identifier and timestamp, with a window size of 60 seconds and a step size of 10 seconds, dividing the data into multiple windows. For the electricity consumption data within each window, real-time change characteristics are calculated. These characteristics include the difference in electricity consumption between the beginning and end of the window, the rate of change per unit time, the fluctuation range of the maximum and minimum values ​​within the window, and the rate of change compared to the previous window. Dynamic thresholds corresponding to each real-time change characteristic are generated according to preset dynamic threshold rules. If any real-time change characteristic in the current window exceeds the corresponding dynamic threshold, it is marked as an abnormal event, generating time-series data containing basic identifiers, time information, real-time change characteristics, and abnormal event markers.

[0018] In electricity monitoring systems, dividing the raw dataset into time windows is fundamental for anomaly detection. For example, a power company analyzes smart meter data from residential areas. Devices are identified by meter IDs, such as "SM-10086," and timestamps are accurate to the second. The system uses a 60-second window size, sliding it every 10 seconds to create overlapping windows. For instance, the first window might be from 08:00:00 to 08:01:00, the second from 08:00:10 to 08:01:10, and so on. In one possible implementation, the real-time change characteristic calculation process is specifically manifested as follows: For device "SM-10086" during the window from 08:00:00 to 08:01:00, the first and last electricity consumptions are 2.5 kWh and 2.8 kWh respectively, with a difference of 0.3 kWh; the rate of change per unit time is 0.3 / 60 = 0.005 kWh / second; the fluctuation range between the maximum value of 2.9 kWh and the minimum value of 2.5 kWh within the window is 0.4 kWh; compared with the last value of 2.4 kWh in the previous window (07:59:50 to 08:00:50), the month-on-month change rate is (2.8-2.4) / 2.4 = 16.7%. It should be noted that the dynamic threshold rules are generated based on the statistical characteristics of historical data. Specifically, the system may calculate the mean and standard deviation of each characteristic based on the electricity consumption data of the same period over the past 30 days, setting the threshold range to the mean plus or minus 3 times the standard deviation; data exceeding this range is marked as abnormal. For example, if the mean of the rate of change per unit time for a certain period is 0.003 kWh / s and the standard deviation is 0.0007 kWh / s, then the upper limit of the dynamic threshold is 0.003 + 3 × 0.0007 = 0.0051 kWh / s, and the lower limit is 0.003 - 3 × 0.0007 = 0.0009 kWh / s. This means that within this period, the rate of change per unit time should normally be between 0.0009 and 0.0051 kWh / s. For the dynamic threshold of the month-on-month change rate, the month-on-month change rate of electricity consumption data for each time window is calculated, along with its mean and standard deviation. The dynamic threshold of the month-on-month change rate is determined by considering the fluctuation characteristics of the electricity consumption data and the tolerance for anomalies in actual business operations. For instance, if the mean of the month-on-month change rate for a certain period is 10% and the standard deviation is 2%, the dynamic threshold is calculated as the mean plus 2.5 times the standard deviation, resulting in a dynamic threshold of 10% + 2.5 × 2% = 15%. In one embodiment, if the rate of change of "SM-10086" within the window from 08:00:00 to 08:01:00 is 0.005 kWh / s and approaches but does not exceed the threshold of 0.0051 kWh / s, the system does not mark it as an anomaly; however, if the month-on-month change rate of 16.7% exceeds the dynamic threshold of 15% for this feature, it is marked as an abnormal event. The generated time-series data includes the basic identifier "SM-10086", the time information "2023-05-15-08:00:00 to 08:01:00", the real-time change feature value, and the "month-on-month change anomaly" label.Preferably, this anomaly detection method based on multiple features and a sliding window can promptly detect abnormal electricity consumption, such as equipment failure or electricity theft, providing decision support for power grid management.

[0019] Step S103: Extract multi-dimensional feature vectors from the time-series data and historical electricity consumption records, and output the classification results of abnormal behavior through a pre-built classifier.

[0020] Time information, real-time change features, and basic identifiers are extracted from the time-series data. Historical electricity consumption records for the past 30 days are obtained using the basic identifiers, including average daily electricity consumption, load curve standard deviation, and average electricity consumption of similar equipment during the same period. The time information, real-time change features, and historical electricity consumption records are concatenated into multi-dimensional features to generate a multi-dimensional feature vector. This multi-dimensional feature vector is input into a pre-built classifier using a random forest algorithm, which outputs classification results for abnormal behavior. The pre-defined classification results for abnormal behavior include normal fluctuation results, suspected equipment failure results, abnormal user behavior results, and abnormal data collection results.

[0021] In electricity monitoring systems, multi-dimensional feature analysis is crucial for identifying abnormal electricity consumption behavior. By extracting time information, real-time change characteristics, and basic identifiers from time-series data, the system can construct a complete electricity consumption profile. For example, the electricity monitoring system of an industrial park collected data on a high-energy-consuming device identified as "DV-10086". The system first extracted the device's basic identifier information, including the device identifier "DV-10086", user number "U2023-567", and meter model "SM-500X". In one possible implementation, the system retrieved the device's historical electricity consumption records for the past 30 days, finding that its average daily electricity consumption was 1250 kWh, and the standard deviation of its load curve was 78.5, while the average electricity consumption of similar devices during the same period was 1180 kWh. This historical data provides a benchmark reference for anomaly detection. Specifically, when the system detects real-time changes within a time window (e.g., a change rate of 25 kW / min, significantly higher than the normal value of 8 kW / min), it concatenates these features with historical electricity consumption records to form a feature vector containing multiple dimensions. It should be noted that the construction process of the multi-dimensional feature vector integrates time-dimensional information (e.g., weekday / non-weekday markings, time period type), real-time changes (e.g., fluctuation amplitude, change rate), and historical electricity consumption records (e.g., average daily electricity consumption, load curve features). Preferably, the pre-trained random forest classifier, after receiving these feature vectors, can output accurate abnormal behavior classification results. For example, when the system detects that equipment suddenly experiences periodic, drastic fluctuations in electricity consumption during normal working hours, and the fluctuation pattern is similar to historical fault patterns, the classifier may output a "suspected equipment fault result," reminding maintenance personnel to check. When the system detects a sudden increase in electricity consumption of a device in the early morning, if step S102 marks it as abnormal, but step S103 finds that the device has historical load fluctuations caused by regular maintenance during the same period, it is classified as a "normal fluctuation result." Understandably, different types of anomalies have their own unique characteristics. For example, "abnormal user behavior results" usually manifest as a sudden and significant increase in electricity consumption without obvious fluctuations, while "abnormal data collection results" may manifest as data jumps or discontinuities.

[0022] Step S104: Obtain the user registration information and the classification results of the abnormal behavior, construct the user-abnormal event association matrix, analyze the user-abnormal event association matrix through the association rule mining algorithm, and generate a multi-source association dataset of abnormal behavior.

[0023] The process involves: acquiring basic attributes from user registration information to generate basic user profile features, including user type, geographic location, registration duration, and device type; reading data corresponding to abnormal behavior classification results to extract core features of abnormal behavior, including anomaly type, occurrence time, duration, and impact range; processing user profile features and core features of abnormal behavior through data cleaning, including removing missing and outliers and normalizing numerical features to obtain standardized user profile features and core features of abnormal behavior; constructing a user-abnormal event association matrix by associating the standardized user profile features and core features of abnormal behavior with user IDs, where each row represents a user and each column represents an abnormal feature; analyzing the user-abnormal event association matrix using the Apriori algorithm, recording rules with confidence levels higher than a preset confidence threshold, and discarding rules with confidence levels lower than the preset confidence threshold; and grouping high-confidence association rules using hierarchical clustering to generate a multi-source association dataset of abnormal behaviors.

[0024] In the business scenario of electricity monitoring systems, generating association rules and constructing multi-source association datasets based on user registration information and abnormal behavior data is an important method for optimizing anomaly detection. For example, obtaining basic attributes from user registration information is the first step in building a user profile. Basic attributes include user type, geographical location, registration duration, and device type. User type can be divided into residential users, commercial users, and industrial users; geographical location refers to the user's location, such as a city or industrial park; registration duration records the length of time the user has been connected to the system, such as 3 years; device type refers to the category of electrical equipment, such as high-energy-consuming equipment or smart home appliances. In one possible implementation, the system extracts information about an industrial user from the database: user type is "industrial user," geographical location is "an industrial park in a certain city," registration duration is "5 years," and device type is "heavy machinery." These attributes constitute the basic user profile features, providing identity basis for subsequent association analysis. Specifically, reading the abnormal behavior classification results and extracting core features is key to abnormal event analysis. Core features include anomaly type, occurrence time, duration, and scope of impact. The anomaly type can be equipment failure, abnormal user behavior, or data collection anomaly; the occurrence time records the specific moment the anomaly occurred, such as "April 20, 2025, 14:00"; the duration refers to the duration of the anomaly, such as "2 hours"; and the scope of impact refers to the affected equipment or area, such as "single device". For example, the system detects an "equipment failure" anomaly on a weekday morning, occurring at "April 21, 2025, 02:00", lasting for 1 hour, affecting only that device. It should be noted that data cleaning ensures the quality of the feature data. The cleaning process includes removing missing and outlier values ​​and normalizing numerical features. Missing values ​​refer to incomplete data records, such as a user's geographic location being empty; outliers refer to data that significantly deviates from the normal range, such as a registration duration recorded as "999 years". Normalization maps numerical features such as registration duration and duration to the range of 0 to 1. For example, if the system finds a user record with a duration of "-1 hour", it identifies it as an outlier and removes it. The registration duration is then normalized, mapping 5 years to 0.5. This cleaned feature data standardization improves the accuracy of subsequent analysis. In one embodiment, constructing a user-abnormal event association matrix requires linking user basic profile features with core abnormal behavior features through user IDs. Each row in the user-abnormal event association matrix represents a user, and each column represents an abnormal feature, such as "duration of equipment failure" or "time of occurrence of the abnormality". For example, the system generates a row of data for user U2023-567, including their industrial user type, 5 years of registration duration, and a 2-hour duration of a certain equipment failure, as shown in the table below:

[0025]

[0026] This matrix structure is clear and facilitates the discovery of associations between users and anomalies. Preferably, the Apriori algorithm is used to analyze the association matrix to uncover high-confidence rules. The Apriori algorithm identifies frequent itemsets and generates rules, such as "Industrial users with registration duration > 3 years → Equipment failure". The confidence threshold is set to 0.8; rules higher than this value are recorded in the association rule database, while those lower are discarded. For example, the system finds that the rule "Commercial users with geographical location in urban areas → Data collection anomaly" has a confidence of 0.85 and retains it. This method effectively filters out reliable anomaly patterns. Hierarchical clustering groups high-confidence rules to generate multi-source association datasets. Hierarchical clustering groups rules based on feature similarity, such as "Equipment failure related rules" or "User behavior anomaly related rules". For example, the system clusters "Industrial users → Equipment failure" and "Registration duration > 5 years → Equipment failure" into one group, generating a multi-source association dataset for equipment failure. This grouping helps the system develop precise response strategies for different anomaly types. Using the methods described above, the system forms a complete chain from user profiling to the processing of abnormal features, and rule mining and grouping further enhance the targeting of anomaly detection. The implementation of each step supports each other, ensuring the reliability and practicality of the analysis results.

[0027] Step S105: Based on the multi-source associated dataset of abnormal behavior, the newly introduced electricity consumption data is processed in real time. Through the incremental update mechanism, if the newly added data triggers the preset abnormal pattern matching conditions, an abnormal behavior warning signal is generated immediately.

[0028] Newly received electricity consumption data is acquired from the real-time data stream. The data stream timestamp and electricity consumption data features are extracted to generate an initial data vector. Using an incremental update mechanism, the initial data vector is compared with the abnormal behavior classification results in a multi-source associated dataset of abnormal behaviors to determine data similarity. If the data similarity is higher than a preset similarity threshold, the abnormal behavior classification result is determined based on a preset abnormal event identifier, resulting in an incremental abnormal behavior classification result. An abnormal behavior early warning signal is generated based on the incremental abnormal behavior classification result using a data processing efficiency optimization method.

[0029] Specifically, in smart grid user behavior monitoring scenarios, real-time data stream processing is the core of abnormal behavior detection. For example, to obtain newly received electricity consumption data from a real-time data stream, key fields in the data packet, such as timestamps and electricity consumption data features, need to be parsed. The timestamp records the specific time the data was generated, such as 2025-04-26-22:30:00, used to track the data's time sequence. Electricity consumption data features include instantaneous power, electricity consumption, and voltage fluctuations. For example, a user's data stream shows an instantaneous power of 2.5 kW and a voltage fluctuation of 5 volts. These features are extracted using a streaming processing framework to form an initial data vector with the structure [timestamp, instantaneous power, voltage fluctuation], such as [2025-04-26-22:30:00, 2.5, 5]. In one possible implementation, an incremental update mechanism is used to compare the initial data vector with the abnormal behavior classification results in a multi-source associated dataset. It should be noted that the multi-source associated dataset contains historical abnormal patterns, such as power surges or voltage anomalies, each pattern corresponding to a feature range. During comparison, the system calculates the similarity between the initial data vector and the abnormal pattern, for example, by measuring Euclidean distance. Assuming a similarity threshold of 0.8, if a data vector has a similarity of 0.85, which is higher than the threshold, it indicates a possible anomaly. Specifically, based on the abnormal event identifier, such as "power surge," the system determines the classification result as abnormal behavior and generates an incremental classification result, such as "User ID_002, power surge, 2025-04-26-22:30:00." Preferably, the data processing efficiency optimization method improves performance through streaming computing and caching mechanisms. For example, the system caches frequently accessed abnormal patterns in memory to reduce database query time. In one embodiment, an abnormal behavior warning signal is generated based on the incremental classification result, and the signal includes the user ID, anomaly type, and timestamp. For example, for user ID_002, an alert signal "Power Surge, 2025-04-26 22:30:00" is generated and sent to the monitoring platform via a message queue. This approach supports real-time response and facilitates rapid intervention. In another embodiment, feature comparison can be supplemented from multiple perspectives. For example, in addition to power and voltage, the timing of electricity consumption can be added; for instance, late-night electricity consumption may be indirect evidence of anomalies. Assuming user ID_002's electricity consumption data shows a sudden surge late at night, combined with historical data analysis, the system further confirms the possibility of an anomaly. This multi-dimensional comparison improves classification accuracy. For example, a user's historical data shows that their nighttime electricity consumption is usually below 0.5 kW, while the current data is 2.5 kW, triggering a higher confidence level for an anomaly determination. It should be noted that after the alert signal is generated, it can be graded according to the severity of the anomaly. For example, a power surge lasting 5 minutes is a low-level anomaly, and lasting 1 hour is a high-level anomaly. The system adjusts the alert priority accordingly; for example, a high-level anomaly is immediately pushed to the administrator.For example, a high-level alert signal is generated for a continuous one-hour power surge in user ID_002, accompanied by detailed anomaly characteristics, facilitating subsequent analysis and processing. This hierarchical mechanism optimizes resource allocation and improves monitoring efficiency.

[0030] Step S106: Use principal component analysis to generate a high-risk event feature vector based on the abnormal behavior warning signal. Integrate the high-risk event feature vector, the original electricity monitoring dataset, and the newly input electricity data to obtain a fused data vector. Input the fused data vector into a convolutional neural network to generate the finally confirmed abnormal behavior data.

[0031] Anomaly behavior warning signals are acquired, and principal component analysis (PCA) is used to generate high-risk event feature vectors based on these signals. The original electricity consumption monitoring dataset and newly received electricity consumption data are acquired, and data cleaning tools are used to process the data, generating a multi-source dataset. A weighted average method is used to integrate the high-risk event feature vectors with the multi-source dataset, resulting in a fused data vector. A convolutional neural network is implemented using the TensorFlow framework, and the fused data vector is input into the convolutional neural network to generate the final confirmed anomaly behavior data. This final confirmed anomaly behavior data includes abnormal electricity consumption records and anomaly types.

[0032] In smart grid user behavior monitoring scenarios, the processing and analysis of abnormal behavior early warning signals are crucial for ensuring electricity safety. For example, after acquiring abnormal behavior early warning signals, high-risk event feature vectors can be generated using Principal Component Analysis (PCA). PCA is a dimensionality reduction technique designed to extract the most representative features from data. Abnormal behavior early warning signals include user information, real-time change characteristics, anomaly types, and association rule information. For instance, a warning signal might contain a user ID, anomaly type such as "power surge," and a timestamp, such as "User ID_003, power surge, 2025-04-27-01:00:00." Through PCA, the system extracts key features from the signal, such as the duration of the anomaly and the magnitude of the power change, generating a feature vector, such as [duration: 30 minutes, power change: 3 kW], for subsequent analysis. In one possible implementation, after acquiring the original electricity monitoring dataset and newly received electricity data, data cleaning tools are needed for processing. The original dataset may contain missing values ​​or noisy data, such as a user's voltage data being empty. Data cleaning tools can fill in missing values ​​or filter outliers using interpolation, such as removing data with voltage values ​​outside a reasonable range. After cleaning, a multi-source dataset is generated, containing features such as timestamps, instantaneous power, and electricity consumption periods. For example, a multi-source dataset records that user ID_003's instantaneous power at 2025-04-27-01:00:00 was 3.2 kilowatts, and the electricity consumption period was late at night, forming a multi-source dataset. Specifically, a weighted average method is used to integrate the feature vectors of high-risk events with the multi-source dataset to generate a fused data vector. The weighted average method assigns weights based on feature importance and normalizes the feature vector of high-risk events and the multi-source dataset. A specific example is as follows: High-risk event feature vector: Assuming it includes "abnormal duration" and "power change amplitude," it is represented as V1 = [duration = 30 minutes, power change = 3 kW], and after normalization, it is V1norm = [0.6, 0.4]; Multi-source dataset features: including "instantaneous power" and "electricity consumption period," represented as V2 = [Instantaneous power = 3.2 kW, electricity consumption period (late night) = 1], normalized to V2norm = [0.5, 0.5]; according to preset weights, the high-risk feature weight is α = [0.7, 0.3], and the multi-source data weight is β = [0.3, 0.7]; fused data vector = α × V1norm + β × V2norm = [0.7 × 0.6 + 0.3 × 0.5, 0.7 × 0.4 + 0.3 × 0.5] = [0.57, 0.43]. This method ensures the effective combination of multi-dimensional information. Preferably, a convolutional neural network is implemented through the TensorFlow framework, and the fused data vector is input into the network to generate the finally confirmed abnormal behavior data. Convolutional neural networks are good at capturing local patterns in data. In one embodiment, the network inputs the fused data vector and outputs abnormal electricity consumption records and abnormality types.For example, for user ID_003, network analysis confirmed the abnormal record as "2025-04-27-01:00:00, power surge of 3.2 kW", with the anomaly type being "abnormal nighttime electricity consumption". Understandably, the network uses multi-layer convolution and pooling to extract deep features, improving classification accuracy. It should be noted that multi-faceted analysis can further support the results. For instance, combining the user's historical electricity consumption habits, the system found that user ID_003's nighttime electricity consumption is usually below 0.5 kW, and the current 3.2 kW is significantly higher, increasing the anomaly confidence. Furthermore, anomaly duration analysis shows that the 30-minute surge conforms to a medium-level anomaly pattern. These aspects collectively confirm the reliability of the abnormal behavior. For example, if the anomaly type is "abnormal nighttime electricity consumption", the system can generate detailed records, including specific time and features, facilitating subsequent tracking. This combination of multi-dimensional analysis and deep learning can effectively identify abnormal behavior in complex scenarios.

[0033] Step S107: Based on the finally confirmed abnormal behavior data, a real-time monitoring dashboard is generated using visualization technology to dynamically display the abnormal behavior of the power supply company's electricity consumption inspection.

[0034] From the finally confirmed abnormal behavior data, abnormal electricity usage records confirmed within the last 30 days are extracted and categorized and labeled according to user type, geographical location, and abnormality type to obtain a structured abnormality dataset. Key performance indicators (KPIs) are calculated based on this structured abnormality dataset, and the sliding window algorithm is used to update the KPI values ​​in real time. These KPIs include anomaly detection accuracy index, warning coverage percentage, and average response time. A multi-level interactive dashboard is constructed using the D3.js visualization library, setting indicator threshold warning lines. If an indicator value exceeds the warning line, a visual alarm signal is triggered, and a heat map of the abnormal area is displayed. A time-series analysis model is established for the dashboard data, using exponential smoothing to predict the changing trends of each indicator over the next 24 hours, generating trend prediction curves. An abnormal event distribution map is drawn using a geographic information system (GIS), automatically clustering anomalies based on density to form high-risk area markers and displaying recommendations for the allocation of inspection resources in each area. WebSocket technology is used to achieve real-time data push; when a new abnormal event is confirmed, the dashboard indicators and charts are automatically updated to maintain the real-time nature of the monitoring data.

[0035] In smart grid user behavior monitoring scenarios, extracting and classifying confirmed abnormal electricity consumption records from the last 30 days of the finally confirmed abnormal behavior data is a key step in optimizing the monitoring system. For example, the system filters records from the abnormal behavior data with a time range of March 28, 2025 to April 26, 2025, including user ID, anomaly type, timestamp, and geographical location. User types can be divided into residential, commercial, and industrial users; geographical locations are divided by administrative region, such as "Chaoyang District, Beijing"; anomaly types include "power surge" and "voltage anomaly." In one possible implementation, the system extracts records through database queries and uses a classification algorithm to generate a structured anomaly dataset according to the above dimensions. For example, if residential user ID_005 experiences a power surge in Chaoyang District at 23:00:00 on April 25, 2025, it is marked as "Resident - Chaoyang District - Power Surge." Specifically, key performance indicators (KPIs) are calculated based on the structured anomaly dataset, and a sliding window algorithm is used to update the KPI values ​​in real time. Key performance indicators (KPIs) include anomaly detection accuracy index, early warning coverage percentage, and average response time. The sliding window algorithm uses a 24-hour window, sliding once per hour to count anomalies within the window. Preferably, the anomaly detection accuracy index is calculated by comparing confirmed anomalies with actual audit results; for example, if 100 anomalies are confirmed within a window and 90 are verified correctly by audit, the accuracy is 90%. The early warning coverage percentage measures the proportion of anomalies covered by the early warning signal; for example, if 80 early warnings are issued on a day, covering 70 actual anomalies, the coverage rate is 87.5%. The average response time records the time from early warning issuance to audit response; for example, an average of 15 minutes. In one embodiment, a multi-layered interactive dashboard is constructed using the D3.js visualization library to display the above indicators and set threshold early warning lines. The dashboard includes bar charts, line charts, and heatmaps to visually present indicator changes. For example, if the accuracy threshold is set to 85%, and the accuracy drops to 80% in a certain hour, the dashboard triggers a red visual alarm signal, and the heatmap highlights high-incidence areas such as Chaoyang District. It should be noted that the heatmap is dynamically colored based on the density of abnormal events, with high-density areas displayed in dark red for easy location of problem areas. Understandably, a time-series analysis model is built based on the dashboard data, using exponential smoothing to predict the trend of indicator changes over the next 24 hours. Exponential smoothing predicts future trends based on a weighted average of historical indicator values. For example, the system analyzes accuracy data from the past 7 days and predicts that accuracy may drop from 90% to 88% in the next 24 hours, generating a trend prediction curve to help managers adjust strategies in advance. In one embodiment, an abnormal event distribution map is created using a geographic information system. The system generates the distribution map based on the latitude and longitude data of abnormal events and identifies high-risk areas using a density clustering algorithm. For example, a residential area in Chaoyang District has the highest density of abnormal events and is marked as a high-risk area; the system suggests increasing the number of inspectors to 5.For example, WebSocket technology can be used to push data in real time, ensuring that the dashboard is updated immediately after a new anomaly is confirmed. Suppose a new anomaly record for a commercial user is added at 01:00:00 on April 26, 2025. WebSocket will push the data to the front end, and the anomaly count and heatmap on the dashboard will be updated in real time, with the accuracy index updated to 89%. This real-time capability ensures the monitoring system's rapid response to anomalies, improving the efficiency of power grid safety management.

[0036] Reference Figure 2 This application also provides a computer device, which may be a server, and its internal structure may be as follows: Figure 2 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor is designed to provide computing and control capabilities. The memory of the computer device includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database of the computer device is used to store data such as raw datasets of electricity consumption monitoring and classification results of abnormal behaviors. The network interface of the computer device is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the personalized recommendation method based on real-time user behavior of any of the above embodiments.

[0037] Those skilled in the art will understand that Figure 2 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer equipment on which the present application is applied.

[0038] This invention also provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the intelligent electricity consumption behavior inspection method based on multi-source data of any of the above embodiments. It is understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.

[0039] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in this application and in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0040] The above description is merely a specific implementation of this specification. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the scope of protection of this specification is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this specification, and these modifications or substitutions should all be covered within the scope of protection of this specification.

Claims

1. A method for intelligent electricity consumption behavior inspection based on multi-source data, characterized in that, The method includes: A distributed computing framework is employed to process real-time electricity consumption data streams collected by sensors and smart meters. Parallel computing technology is used to integrate and standardize the electricity consumption data streams, generating a raw electricity consumption monitoring dataset. For the dynamically updated raw electricity consumption monitoring dataset, a sliding window algorithm is applied to extract electricity consumption data within a specified time period, calculate real-time change characteristics, and if the detected real-time change characteristics exceed the corresponding dynamic threshold, they are marked as abnormal events, generating time-series data. The time-series data includes basic identifiers, time information, real-time change characteristics, and abnormal markers. The basic identifiers include device identifier, user number, and meter model. Multi-dimensional feature vectors are extracted from the time-series data and historical electricity consumption records, and a pre-built classifier is used to output the classification results of abnormal behavior. User registration information and the classification results of the abnormal behavior are obtained to construct... The user-abnormal event association matrix is ​​analyzed using association rule mining algorithms to generate a multi-source association dataset of abnormal behaviors. Based on this dataset, newly input electricity consumption data is processed in real-time. Through an incremental update mechanism, if new data triggers preset abnormal pattern matching conditions, an abnormal behavior warning signal is generated immediately. Principal component analysis is used to generate high-risk event feature vectors based on the warning signals. These vectors, along with the original electricity consumption monitoring dataset and the newly input electricity consumption data, are integrated to obtain a fused data vector. This fused data vector is then input into a convolutional neural network to generate the final confirmed abnormal behavior data. Based on this data, a real-time monitoring dashboard is generated using visualization technology to dynamically display the abnormal behaviors identified during the power supply company's electricity consumption audits.

2. The method according to claim 1, characterized in that, The system employs a distributed computing framework to process real-time electricity consumption data streams collected by sensors and smart meters. Parallel computing technology is used to integrate and standardize the electricity consumption data streams, generating a raw dataset for electricity monitoring, including: The system acquires real-time electricity consumption data streams from sensors and smart meters using a distributed computing framework, and then uses parallel computing technology to segment the data streams to obtain preliminarily processed data segments. Extract multi-source data from the initially processed data segments, and use a distributed stream processing platform to unify the format of the multi-source data to obtain an integrated data set. If the field formats are inconsistent in the integrated dataset, Pandas is used to adjust the field types and units to obtain the original electricity monitoring dataset.

3. The method according to claim 1, characterized in that, For the dynamically updated raw dataset of electricity consumption monitoring, a sliding window algorithm is applied to extract electricity consumption data within a specified time period, calculate real-time change characteristics, and if the detected real-time change characteristics exceed the corresponding dynamic threshold, it is marked as an abnormal event, generating time-series data. The time-series data includes a basic identifier, time information, real-time change characteristics, and anomaly markers. The basic identifier includes the device identifier, user number, and meter model, including: The dynamically updated raw dataset of electricity consumption monitoring is grouped by device identifier and timestamp, with a window size of 60 seconds and a step size of 10 seconds, and divided into multiple windows. For the electricity consumption data in each window, calculate the real-time change characteristics, which include the difference in electricity consumption at the beginning and end of the window, the rate of change per unit time, the fluctuation range of the maximum and minimum values ​​within the window, and the month-on-month change rate compared with the previous window. Based on preset dynamic threshold rules, dynamic thresholds corresponding to each real-time change feature are generated. If any real-time change feature in the current window exceeds the corresponding dynamic threshold, it is marked as an abnormal event, and time-series data containing basic identifiers, time information, real-time change features, and abnormal markers is generated.

4. The method according to claim 1, characterized in that, The step involves extracting multi-dimensional feature vectors from the time-series data and historical electricity consumption records, and then outputting classification results of abnormal behavior through a pre-built classifier, including: Extract time information, real-time change features, and basic identifiers from the time-series data; The historical electricity consumption records of the equipment over the past 30 days are obtained through basic identification. These historical electricity consumption records include the average daily electricity consumption, the standard deviation of the load curve, and the average electricity consumption of similar equipment during the same period. Time information, real-time change characteristics, and historical electricity consumption records are combined into multi-dimensional features to generate a multi-dimensional feature vector. The multi-dimensional feature vectors are input into a pre-built classifier using the random forest algorithm, and the output is the classification result of abnormal behavior. The pre-defined classification results of abnormal behavior include normal fluctuation results, suspected equipment failure results, abnormal user behavior results, and abnormal data collection results.

5. The method according to claim 1, characterized in that, The process involves obtaining user registration information and the classification results of the abnormal behavior, constructing a user-abnormal event association matrix, analyzing the user-abnormal event association matrix using an association rule mining algorithm, and generating a multi-source association dataset of abnormal behavior, including: Obtain basic attributes from user registration information to generate basic user profile features. These basic attributes include user type, geographic location, registration duration, and device type. Read the data corresponding to the abnormal behavior classification results and extract the core features of the abnormal behavior. The core features include the abnormal type, occurrence time, duration and scope of impact. The data cleaning process processes the basic user profile features and core abnormal behavior features, including removing missing and outlier values ​​and normalizing numerical features to obtain standardized basic user profile features and core abnormal behavior features. Based on the user ID associated with standardized user profile features and core features of abnormal behavior, a user-abnormal event association matrix is ​​constructed. In this matrix, each row represents a user and each column represents an abnormal feature. The Apriori algorithm is used to analyze the user-abnormal event association matrix. If the confidence of an association rule is higher than the preset confidence threshold, the rule is recorded in the association rule base. If the confidence of an association rule is lower than the preset confidence threshold, the rule is discarded. High-confidence association rules are grouped using hierarchical clustering methods to generate a multi-source association dataset of anomalous behavior.

6. The method according to claim 1, characterized in that, The multi-source correlation dataset based on abnormal behavior performs real-time stream processing on newly received electricity consumption data. Through an incremental update mechanism, if new data triggers preset abnormal pattern matching conditions, an abnormal behavior early warning signal is generated immediately, including: Newly received electricity consumption data is obtained from the real-time data stream. The data stream timestamp and electricity consumption data features are extracted to generate an initial data vector. By using an incremental update mechanism, the initial data vector is compared with the abnormal behavior classification results in the multi-source association dataset of abnormal behavior to determine the data similarity. If the data similarity is higher than the preset similarity threshold, the abnormal behavior classification result is determined according to the preset abnormal event identifier, and the incremental abnormal behavior classification result is obtained. An abnormal behavior warning signal is generated based on the incremental abnormal behavior classification results by employing a data processing efficiency optimization method.

7. The method according to claim 1, characterized in that, The process involves using principal component analysis to generate high-risk event feature vectors based on abnormal behavior warning signals. These feature vectors, along with the original electricity monitoring dataset and newly input electricity data, are integrated to obtain a fused data vector. This fused data vector is then input into a convolutional neural network to generate the final confirmed abnormal behavior data, including: Obtain abnormal behavior warning signals and use principal component analysis to generate high-risk event feature vectors based on the abnormal behavior warning signals; Obtain the original electricity consumption monitoring dataset and the newly input electricity consumption data, and use data cleaning tools to process the data to generate a multi-source data set; A weighted average method is used to integrate the feature vectors of high-risk events with multi-source datasets to obtain a fused data vector; A convolutional neural network is implemented using the TensorFlow framework. The fused data vector is input into the convolutional neural network to generate the final confirmed abnormal behavior data, which includes abnormal electricity consumption records and abnormal types.

8. The method according to claim 1, characterized in that, Based on the finally confirmed abnormal behavior data, a real-time monitoring dashboard is generated using visualization technology to dynamically display the abnormal behaviors detected by the power supply company's electricity consumption audit, including: Extract the confirmed abnormal electricity usage records from the last 30 days from the finally confirmed abnormal behavior data, and classify and label them according to user type, geographical location and abnormality type to obtain a structured abnormal dataset. Key performance indicators (KPIs) are calculated based on a structured anomaly dataset, and the KPI values ​​are updated in real time using a sliding window algorithm. The KPIs include anomaly detection accuracy index, early warning coverage percentage, and average response time. A multi-layered interactive dashboard is built using the D3.js visualization library. Indicator threshold warning lines are set. If the indicator value exceeds the warning line, a visual alarm signal is triggered, and a heat map of the abnormal area is displayed. A time-series analysis model is established for dashboard data, and the exponential smoothing method is used to predict the changing trends of various indicators in the next 24 hours, generating a trend prediction curve. By combining geographic information systems to draw distribution maps of abnormal events, high-risk areas are automatically clustered based on anomaly density, and suggestions for the allocation of inspection resources in each area are displayed. Using WebSocket technology, data is pushed in real time. When a new abnormal event is confirmed, the dashboard indicators and charts are automatically updated to keep the monitoring data real-time.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Group behavior power distribution network power consumption abnormity detection method based on user attribute tags

    CN110288383A

  • Anti-electricity-stealing inspection monitoring method and anti-electricity-stealing inspection monitoring system

    CN113239087A