Fog computing-based loss detection method for smart grids and system therefor
The hybrid NTL detection method using ARIMA and random forest algorithms addresses inefficiencies in smart grid fraud detection, improving accuracy and reducing costs while preventing overfitting, thus ensuring energy sustainability.
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- IND ACADEMIC COOPERATION FOUND UNIV OF INCHEON
- Filing Date
- 2025-04-29
- Publication Date
- 2026-07-21
AI Technical Summary
Existing smart grid systems face challenges in detecting non-technical losses (NTL) due to energy theft, meter tampering, and billing errors, with limitations such as high computational costs, overfitting, and inefficient detection methods, leading to significant financial losses and sustainability issues.
A hybrid non-technical-loss detection (HNTLD) method using ARIMA forecasting and random forest machine learning to identify fraudulent smart meters by comparing smart meter readings with observer meter data, with threshold checks and approximate reading predictions to enhance detection accuracy and reduce computational costs.
The method improves detection accuracy, reduces execution time, and prevents overfitting, effectively identifying fraudulent smart meters and reducing computational costs, thereby enhancing energy sustainability and financial stability in smart grids.
Smart Images

Figure 112025048641037-PAT00040_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a method for detecting loss in a fog computing-based smart grid and a system for the same. Background Technology
[0002] The primary goal of the smart grid (SG) is to enhance the energy efficiency, sustainability, resilience, security, and dependability of existing power grids. Energy sustainability is a critical factor in the growth of anthropogenic activities, communities, and civilization. As electricity plays a vital role in the residential, commercial, transportation, and industrial sectors, it will continue to be a key factor influencing socioeconomic changes in all nations. The smart grid can preserve the environment and end energy crises by intelligently and dynamically managing distributed energy resource assets to reap various socioeconomic and environmental benefits.
[0003] The latest Information and Communications Technology (ICT) has been integrated into smart grids to improve two-way communication between energy producers and consumers. A key component of the smart grid environment is the Advanced Metering Infrastructure (AMI), which provides state-of-the-art electronic software and hardware systems for automatic pricing, billing infrastructure, and metering. Smart meters (SM) are installed in each home or business to record energy consumption, which is then recorded by the Advanced Metering Infrastructure.
[0004] The sustainability of energy resources is related to a series of requirements, including high efficiency, low environmental and ecological impact, sustainable energy resources, other non-technical aspects of energy sustainability, and economic sustainability and economic feasibility. Non-technical loss (NTL) in smart grids refers to energy that has been distributed but not billed, primarily due to energy theft. Energy fraud can lead to the unequal distribution of costs, and honest consumers end up subsidizing the illicit gains of fraudsters.
[0005] According to a recent study focusing on the financial sustainability of electricity, developed countries suffer annual losses of over $6 billion due to smart grid theft (SG theft), whereas India suffers annual losses of $16.2 billion due to energy theft. The study also found that transmission and distribution losses reach up to 12.5% in South Korea, 9% in Taiwan, and 7% in Japan. High NTLs, estimated at double-digit percentages, are threatening the financial sustainability of energy utilities in several countries.
[0006] Theft is the primary cause of most of these losses. NTL fraud involves manipulating smart meters and infiltrating networks to illegally profit from consumers. This silent crime disrupts the normal billing procedures of power utilities, resulting in revenue losses and wasting people's energy-saving efforts. Energy theft has a negative financial impact on businesses and can result in unreasonable additional costs for consumers who pay their bills on time.
[0007] Energy fraud leads to excessive energy consumption, which contradicts sustainability goals. The supply of sufficient, reliable, and affordable energy that complies with social and environmental norms is a critical element of energy sustainability. Experts and researchers from smart grid academia and industry have considered various approaches to effectively address the NTL problem.
[0008] In many countries, the conventional approach to detecting energy fraud is a manual method that involves manually analyzing notable trends and patterns in consumer power consumption data and physically inspecting and verifying suspicious smart meters. This solution is not cost-effective because conducting numerous field investigations and detecting fraudulent activities incurs significant operational costs.
[0009] Smart meters are vulnerable to energy theft, which can occur in various ways, such as direct wire hooking or tapping (to cause reading errors and low effective consumption), root-level firmware access, or data corruption during transmission, recording, or storage.
[0010] According to non-patent literature [8], one of the most recent and important solutions to the NTL identification problem is the use of smart meter data. Machine learning and deep learning models are applied to analyze readings captured from smart meters and evaluate usage patterns to detect fraudulent activity. Machine learning techniques such as decision trees (non-patent literature
[10] ), neural networks (non-patent literature
[11] ), support vector machines (non-patent literature
[12] ), ARIMA (auto-regressive integrated moving average) modeling for verification (non-patent literature
[13] ), random forests (non-patent literature 14)), and fuzzy models (non-patent literature
[15] ) are utilized to identify suspicious consumer data in supervised and unsupervised classification systems.
[0011] However, existing NTL detection systems for smart grids have limitations such as execution time, overfitting, class imbalance, and high computational costs. Due to these limitations, there is a growing demand for a more robust and efficient solution for detecting NTL fraud in smart grids, regardless of whether observer meters or observer smart meters (ObSMs) are installed. Prior art literature
[0012] [1] Adil M, Javaid N, Qasim U, Ullah I, Shafiq M, Choi J-G. Lstm and bat-based rusboost approach for electricity theft detection. Appl Sci 2020;10(12):4378.[2] Rosen MA. Energy sustainability with a focus on environmental perspectives. Earth Syst Environ 2021;5(2):217-30.[3] Singh SK, Bose R, Joshi A. Energy theft detection in advanced metering infrastructure. In: 2018 IEEE 4th world forum on internet of things. IEEE; 2018, p. 529-34.[4] Arora M. Power transmission and distribution losses in india-a study report. J Curr Sci 2019;20(1).[5] Avila NF, Figueroa G, Chu C-C. NTL detection in electric distribution systems using the maximal overlap discrete wavelet-packet transform and random undersampling boosting. IEEE Trans Power Syst 2018;33(6):7171-80.[6] Nazari-Heris M, Mirzaei MA, Mohammadi-Ivatloo B, Marzband M, Asadi S. Economic-environmental effect of power to gas technology in coupled electricity and gas systems with price-responsive shiftable loads. J Clean Prod 2020;244:118769.[7] Sathe MT, Adamuthe AC. Comparative study of supervised algorithms for prediction of students’ performance. Int J Mod Educ Comput Sci 2021;13(1).[8] Rengaraju P, Pandian SR, Lung C-H. Communication networks and non-technical energy loss control system for smart grid networks. In: 2014 IEEE innovative smart grid technologies-Asia. IEEE; 2014, p. 418-23.[9] Ahmed M, Khan A, Ahmed M, Tahir M, Jeon G, Fortino G, Piccialli F. Energy theft detection in smart grids: taxonomy, comparative analysis, challenges, and future research directions. IEEE / CAA J Autom Sin 2022;9(4):578-600.
[10] Salman Saeed M, Mustafa MW, Sheikh UU, Jumani TA, Khan I, Atawneh S, et al. An efficient boosted c5. 0 decision-tree-based classification approach for detecting non-technical losses in power utilities. Energies 2020;13(12):3242.
[11] Pereira J, Saraiva F. Convolutional neural network applied to detect electricity theft: A comparative study on unbalanced data handling techniques.Int J Electr Power Energy Syst 2021;131:107085.
[12] Haq EU, Huang J, Xu H, Li K, Ahmad F. A hybrid approach based on deep learning and support vector machine for the detection of electricity theft in power grids. Energy Rep 2021;7:349-56.
[13] Box GE, Jenkins GM, Reinsel GC, Ljung GM. Time series analysis: forecasting and control. John Wiley & Sons; 2015.
[14] Abdulkareem NM, Abdulazeez AM, et al. Machine learning classification based on Radom Forest Algorithm: A review. Int J Sci Bus 2021;5(2):128-42.
[15] Jaiswal S, Ballal MS. Fuzzy inference based electricity theft prevention system to restrict direct tapping over distribution line. J Electr Eng Technol 2020;15:1095-106.
[16] Khan ZA, Adil M, Javaid N, Saqib MN, Shafiq M, Choi J-G. Electricity theft detection using supervised learning techniques on smart meter data. Sustainability 2020;12(19):8023.
[17] Mujeeb S, Javaid N, Khalid R, Imran M, Naseer N.DE-RUSBoost: An efficient electricity theft detection scheme with additive communication layer. In: ICC 2020-2020 IEEE international conference on communications. IEEE; 2020, p.1-6.
[18] Ullah A, Javaid N, Samuel O, Imran M, Shoaib M. CNN and GRU based deep neural network for electricity theft detection to secure smart grid. In: 2020 international wireless communications and mobile computing. IEEE; 2020, p.1598-602.
[19] Yao D, Wen M, Liang X, Fu Z, Zhang K, Yang B. Energy theft detection with energy privacy preservation in the smart grid. IEEE Internet Things J 2019;6(5):7659-69.
[20] Zheng Z, Yang Y, Niu X, Dai H-N, Zhou Y. Wide and deep convolutional neural networks for electricity-theft detection to secure smart grids. IEEE Trans Ind Inf 2017;14(4):1606-15.
[21] Han W, Xiao Y. FNFD: a fast scheme to detect and verify non-technical loss fraud in smart grid. In: Proceedings of the 2016 ACM international on workshop on traffic measurements for cybersecurity. 2016, p. 24-34.
[22] Jamil A, Alghamdi TA, Khan ZA, Javaid S, Haseeb A, Wadud Z, et al. An innovative home energy management model with coordination among appliances using game theory. Sustainability 2019;11(22):6287.
[23] Lee J, Sun YG, Sim I, Kim SH, Kim DI, Kim JY. Non-technical loss detection using deep reinforcement learning for feature cost efficiency and imbalanced dataset. IEEE Access 2022;10:27084-95.
[24] Kumar N, Aujla GS, Das AK, Conti M. ECCAuth: a secure authentication protocol for demand response management in a smart grid system. IEEE Trans Ind Inf 2019;15(12):6572-82.
[25] Costa BC, Alberto BL, Portela AM, Maduro W, Eler EO. Fraud detection in electric power distribution networks using an ann-based knowledge-discovery process. Int J Artif Intell Appl 2013;4(6):17.
[26] Krishna VB, Iyer RK, Sanders WH. ARIMA-based modeling and validation of consumption readings in power grids. In: International conference on critical information infrastructures security. Springer; 2015, p. 199-210.
[27] Han W, Xiao Y. CNFD: a novel scheme to detect colluded non-technical loss fraud in smart grid. In: International conference on wireless algorithms, systems, and applications. Springer; 2016, p. 47-55.
[28] Khan HM, Khan A, Jabeen F, Rahman AU. Privacy preserving data aggregation with fault tolerance in fog-enabled smart grids. Sustainable Cities Soc 2021;64:102522.
[29] El-Sayed OH, Emam O, Abdel-Salam M. Deep learning framework for physical internet hubs inbound containers forecasting. Int J Adv Comput Sci Appl 2022;13(3).
[30] Comden J, Zamzam AS, Bernstein A. Data-driven chance-constrained design of voltage droop control for distribution networks. Tech. rep, Golden, CO (United States): National Renewable Energy Lab.(NREL); 2022.
[31] State grid corporation of China. 2021, http: / / www.sgcc.com.cn. [Accessed 3 January 2021].
[32] Badawi SA, Guessoum D, Elbadawi I, Albadawi A.A novel time-series transformation and machine-learning-based method for ntl fraud detection in utility companies. Mathematics 2022;10(11):1878.
[33] Idris A, Rizwan M, Khan A. Churn prediction in telecom using Random Forest and PSO based data balancing in combination with various feature selection strategies. Comput Electr Eng 2012;38(6):1808-19.
[34] Zidi S, Mihoub A, Qaisar SM, Krichen M, Al-Haija QA. Theft detection dataset for benchmarking and machine learning based classification in a smart grid environment. J King Saud Univ Comput Inf Sci 2022.
[35] Rstudio. 2021, https: / / www.rstudio.com / products / rstudio / download / . [Accessed 15 March 2021].
[36] Han W, Xiao Y. NFD: a practical scheme to detect non-technical loss fraud in smart grid. In: 2014 IEEE international conference on communications. IEEE; 2014, p. 605-9.
[37] Individual household electric power consumptiondata set. 2020, https: / / archive.ics.uci.edu / ml / datasets / Individual+household+electric+power+consumption.[Accessed 23 June 2020]. The problem to be solved
[0013] An embodiment of the present invention provides a smart grid loss detection method with high detection accuracy for non-technical losses.
[0014] In addition, embodiments of the present invention provide a method for detecting loss in a smart grid with reduced execution time.
[0015] In addition, embodiments of the present invention provide a method for detecting loss in a smart grid that can prevent overfitting.
[0016] In addition, embodiments of the present invention provide a method for detecting loss in a smart grid that can reduce computational costs. means of solving the problem
[0017] In one aspect, an embodiment of the present invention comprises: (A) a step of checking whether a reading of an observer meter (ObSM) is available; (B) if the reading of the observer meter is available, (B-1) a step of calculating the difference between the reading collected from a plurality of smart meters (SM) and the reading of the observer meter; (B-2) a step of determining whether the smart meter is a fraudulent smart meter based on the difference between the reading of the smart meter and the reading of the observer meter, wherein if the difference between the reading of the smart meter and the reading of the observer meter is less than a predetermined threshold, it is determined to be normal, and if it exceeds the threshold, an individual row in which fraud has occurred is identified; (B-3) after the individual row is identified, a step of separating the dataset of the smart meter into normal data and suspicious data; (B-4) a step of predicting an approximate reading of the suspicious smart meter using ARIMA (auto-regressive integrated moving average); The present invention provides a method for detecting loss in a smart grid, comprising: (B-5) a step of identifying a fraudulent smart meter by comparing the predicted value of the ARIMA with the actual reported reading; and (C) if the reading of the observer meter is unavailable, (C-1) a step of predicting an approximate reading of a suspected smart meter using ARIMA (auto-regressive integrated moving average); and (C-2) a step of identifying a fraudulent smart meter by comparing the predicted value of the ARIMA with the actual reported reading.
[0018] An embodiment of the present invention provides a method for detecting loss in a smart grid, wherein the loss is a non-technical loss.
[0019] In an embodiment of the present invention, the step of identifying the individual row in step (B-2) is performed by calculating the following mathematical formula 4 for each row, and go ( is the allowable deviation, and if it is 0.02 or less, it is a normal reading. go A method for detecting loss in a smart grid is provided, wherein if it is less than a fraudulent reading, it is a fraudulent reading.
[0020] [Mathematical Formula 4]
[0021]
[0022] (here, r j,j period T j At SM i It is the electricity reading consumed at, and E j is the total ObSM electrical reading in the j-th period.)
[0023] An embodiment of the present invention provides a method for detecting loss in a smart grid, wherein the threshold value is constructed using Chebyshev's theorem.
[0024] An embodiment of the present invention provides a method for detecting loss in a smart grid, wherein the step of identifying the fraudulent smart meter in steps (B-5) and (C-2) is performed by a random forest machine learning model.
[0025] An embodiment of the present invention provides a smart grid loss detection method in which the step of identifying the fraudulent smart meter in steps (B-5) and (C-2) is performed by the following mathematical formula 5, wherein if the result of the mathematical formula 5 continuously yields true for a predetermined period or longer, the smart meter is normal.
[0026] [Mathematical Formula 5]
[0027]
[0028] (Here, ARIMA_forecast i is the value predicted by the ARIMA model for smart meter i, and Suspected_fraudulent_data i is the actual measurement value for the section suspected of fraud regarding smart meter i.)
[0029] In another aspect, an embodiment of the present invention comprises: a memory; and a processor connected to the memory and configured to execute computer-readable instructions included in the memory, wherein the processor comprises: (A) an operation to check whether a reading of an observer meter (ObSM) is available; (B) if the reading of the observer meter is available, (B-1) an operation to calculate the difference between the reading collected from a plurality of smart meters (SM) and the reading of the observer meter; (B-2) an operation to determine whether the smart meter is a fraudulent smart meter based on the difference between the reading of the smart meter and the reading of the observer meter, wherein if the difference between the reading of the smart meter and the reading of the observer meter is less than a predetermined threshold, it is determined to be normal, and if it exceeds the threshold, an operation to identify an individual row in which fraud has occurred; (B-3) after the individual row is identified, an operation to separate the dataset of the smart meter into normal data and suspicious data; (B-4) an operation of predicting an approximate reading of a suspected smart meter using an auto-regressive integrated moving average (ARIMA); and (B-5) an operation of identifying a fraudulent smart meter by comparing the predicted value of the ARIMA with the actual reported reading; and (C) if the reading of the observer meter is not available, (C-1) an operation of predicting an approximate reading of a suspected smart meter using an auto-regressive integrated moving average (ARIMA); and (C-2) an operation of identifying a fraudulent smart meter by comparing the predicted value of the ARIMA with the actual reported reading; the present invention provides a smart grid loss detection system that can be executed.
[0030] In another aspect, an embodiment of the present invention comprises a program stored in a computer-readable storage medium for executing a loss detection method performed by a loss detection system of a smart grid, comprising: (A) an operation of checking whether a reading of an observer meter (ObSM) is available; (B) if the reading of the observer meter is available, (B-1) an operation of calculating the difference between the reading collected from a plurality of smart meters (SM) and the reading of the observer meter; (B-2) an operation of determining whether the smart meter is a fraudulent smart meter based on the difference between the reading of the smart meter and the reading of the observer meter, wherein if the difference between the reading of the smart meter and the reading of the observer meter is less than a predetermined threshold, it is determined to be normal, and if it exceeds the threshold, an operation of identifying an individual row in which fraud has occurred; (B-3) after the individual row is identified, an operation of separating the dataset of the smart meter into normal data and suspicious data; (B-4) an operation to predict an approximate reading of a suspected smart meter using ARIMA (auto-regressive integrated moving average); and (B-5) an operation to identify a fraudulent smart meter by comparing the predicted value of the ARIMA with the actual reported reading; and (C) if the reading of the observer meter is not available, (C-1) an operation to predict an approximate reading of a suspected smart meter using ARIMA (auto-regressive integrated moving average); and (C-2) an operation to identify a fraudulent smart meter by comparing the predicted value of the ARIMA with the actual reported reading; provided a program stored on a computer-readable storage medium. Effects of the invention
[0031] According to an embodiment of the present invention, the detection accuracy for non-technical losses can be improved.
[0032] In addition, according to an embodiment of the present invention, the execution time of the detection method can be reduced.
[0033] In addition, according to an embodiment of the present invention, overfitting can be prevented.
[0034] In addition, according to an embodiment of the present invention, computation costs can be reduced. Brief explanation of the drawing
[0035] FIG. 1 is a flowchart illustrating an NTL fraud detection method according to an embodiment of the present invention. FIG. 2 is a diagram showing a random forest-based HNTLDS model according to an embodiment of the present invention. Figure 3 is a diagram showing fraud and non-fraud SM. Figure 4 is a diagram illustrating a method for detecting SM fraud and non-fraud behavior. FIG. 5 is a diagram showing a comparison between a system according to an embodiment of the present invention and a conventional system. FIG. 6 is another drawing showing a comparison between a system according to an embodiment of the present invention and a conventional system. FIG. 7 is another drawing showing a comparison between a system according to an embodiment of the present invention and a conventional system. FIG. 8 is a block diagram showing a loss detection system for a fog computing-based smart grid according to an embodiment of the present invention. Specific details for implementing the invention
[0036] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Identical or similar components are assigned the same reference number regardless of drawing symbols, and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably solely for the ease of drafting the specification and do not inherently possess distinct meanings or roles.
[0037] In this description, expressions such as “include,” “equip,” or “compose” are intended to refer to certain characteristics, numbers, steps, actions, elements, parts or combinations thereof, and should not be interpreted to exclude the existence or possibility of one or more other characteristics, numbers, steps, actions, elements, parts or combinations thereof other than those described.
[0038] In addition, when describing the embodiments disclosed in this specification, if it is determined that a detailed description of related prior art could obscure the essence of the embodiments disclosed in this specification, such detailed description is omitted.
[0039] The attached drawings are intended only to facilitate understanding of the embodiments disclosed in this specification, and the technical concept disclosed in this specification is not limited by the attached drawings; it should be understood that all modifications, equivalents, and substitutions included within the concept and technical scope of the present invention are included.
[0040] In addition, the disclosures of the papers cited throughout this specification are incorporated by reference in their entirety to more clearly explain the level of the art to which this application pertains and the content of this application.
[0041] One of the critical challenges of smart grid systems is non-technical loss (NTL) in electricity consumption. These losses are primarily caused by energy theft, meter tampering, unauthorized connections, and billing errors. Early detection of these abnormal consumption patterns is crucial for protecting the revenue of power companies and maintaining the stability and efficiency of energy infrastructure.
[0042] Embodiments of the present invention aim to contribute to the achievement of energy sustainability by providing a method or system for detecting energy fraud. Requirements for energy sustainability include economic sustainability and economic feasibility. One of the major factors affecting economic feasibility and economic sustainability is energy fraud.
[0043] Embodiments of the present invention provide two hybrid non-technical-loss detection (HNTLD) schemes for detecting energy fraud in a fog computing-based smart grid. Specifically, they provide an HNTLD that supports observer smart meters (HNTLDS-ObSM) and an HNTLD that does not support observer smart meters (HNTLDS-WObSM). The method according to the embodiments of the present invention considers ARIMA and machine learning techniques to identify energy fraud.
[0044] In addition, extensive simulations were performed on actual power consumption data sets to evaluate the efficiency of the system according to the embodiment of the present invention. The results of the system according to the embodiment of the present invention were compared with existing NTL fraud detection methods. To verify the performance of the system according to the embodiment of the present invention, benchmark functions of various dimensions, such as recall, F1 score, accuracy, sensitivity, precision, and AUC (area under the curve), were used.
[0045] Various methodologies have been used to identify electricity theft, including network-based methods (non-patent literature
[21] ), game theory-based methods (non-patent literature
[22] ), machine learning-based methods (non-patent literature
[23] ), and other methods (non-patent literature
[24] ). In non-patent literature [5], MODWPT and RUSBoost were used to detect NTL. The RUSBoost method is optimal for processing distorted datasets. However, it fails to optimize the classification process and reduces the dataset size, leading to underfitting of the model.
[0046] Non-patent literature
[19] proposes an energy theft detection system using a convolutional neural network (CNN) and Paillier encryption, but it has the problem of high communication and computation costs. Non-patent literature
[25] uses an artificial neural network (ANN) for NTL detection. In this method, a database is created from customer energy data, and then the ANN method is applied to detect fraud cases. A new method based on highly imbalanced labeling data (non-patent literature
[20] ), the wide and deep CNN (WADCNN), is used to detect electricity fraud.
[0047] Non-patent literature
[18] proposed an HDNN technique for detecting NTL using CNN and GRU-PSO. The performance of the proposed technique is measured using standard performance metrics such as accuracy, F1 score, precision, recall, and AUC. This method is powerful and efficient, but it produces results with reduced accuracy on various datasets due to overfitting. Non-patent literature
[16] detects energy theft using the Adasyn algorithm, which uses metaheuristics and deep neural networks. It detects fraudulent customers by submitting consumption data to a visual geometry group (VGG-16).
[0048] FA-XGBoost (Firefly algorithm-based extreme gradient-boosting) is used to classify data, but the number of runs increases as the dataset size increases. Non-patent literature
[26] proposes an ARIMA prediction method to verify smart meter readings, but does not discuss the impact of reduced customer usage due to seasonality. Non-patent literature
[21] and
[27] propose a mathematical model to detect NTL fraud by generating a functional relationship between power consumption data and reported power data. However, they did not consider seasonality and other criteria for distinguishing legitimate users.
[0049] An embodiment of the present invention provides a novel NTL fraud detection approach that uses ARIMA prediction as a predictor to learn the strong correlation between previous data behavior and observer meter readings when observer smart meter or observer meter data is available in a smart grid.
[0050] An observer meter or observer smart meter may be a smart meter installed for the purpose of monitoring and surveillance. In an embodiment of the present invention, NTL fraud is identified using ARIMA and random forest machine learning techniques. Due to the efficiency and simplicity of the hybrid model according to an embodiment of the present invention, a smart grid for fraud detection can be easily implemented.
[0052] System Model and Security Objectives
[0053] Below, the system model, the attacker model, and the performance and security objectives of the method according to an embodiment of the present invention are described.
[0055] System Model
[0056] The system model consists of the following entities.
[0057] 1. Trusted Authority (TA): The TA is responsible for generating keys for all entities. Therefore, when a smart meter withdraws from or joins a grid network, it updates the corresponding information in the TA database.
[0058] 2. Cloud-Control-Center (CCC): The CCC is a trusted entity in the smart grid. The CCC receives metering data to manage smart grid operations such as demand response, forecasting, invoicing, NTL detection, and analysis.
[0059] 3. Fog Node (FN): The FN acts as an intermediary node between smart meters and the CCC, collecting smart meter data and submitting it to the CCC. In a distributed architecture, it can also be used to process data from smart meters in specific regions if necessary to detect energy fraud.
[0060] 4. Smart Meter (SM): The SM is an AMI device installed at the customer's premises that is responsible for submitting usage data to the CCC via the FN. Data collected by the SM is used to develop energy system operation models, which can improve decision-making and increase efficiency. Additionally, it can be used to identify trends in energy production and consumption to provide information to policymakers.
[0061] 5. Observer Smart Meter or Observer Meter (ObSM): ObSM tracks the total energy used by an SM group over a specific period. Fraud is detected by comparing SM and ObSM data.
[0063] A utility company collects metering values or readings from an SM to identify and detect abnormal power usage in order to use the energy fraud detection technology according to an embodiment of the present invention. To store the readings immediately, a secure and reliable high-capacity data storage system, either on-premise or cloud-based, is required.
[0065] attacker model
[0066] The method according to an embodiment of the present invention assumes that CCC and FN are curious and honest. That is, they will try to look at the contents of the metering data. SM is generally susceptible to tampering. However, customers may need to submit accurate data to reduce their electricity bills.
[0067] The method according to an embodiment of the present invention considers the following security attacks and assumptions.
[0068] 1. CCC and FN are considered honest, but they are curious and try to look at the contents of SM data.
[0069] 2. It is assumed that SM is not honest.
[0070] 3. Communication channels may not be safe.
[0071] 4. An internal adversary who knows the SM configuration can manipulate SM data.
[0072] 5. SM may fail to transmit data due to a malfunction.
[0073] 6. An adversary may initiate a False Data Injection (FDI) attack to alter metering data during transmission.
[0074] 7. An attacker can control an adjacent SM and submit inaccurate data to reduce electricity costs.
[0076] Security objectives
[0077] The system according to an embodiment of the present invention achieves the following security objectives.
[0078] 1. Resistance to FDI attacks: If metering data is modified during transit through SM data tempering or FDI attacks, it can be filtered out.
[0079] 2. Integrity: Protects the integrity of metering data and identifies unauthorized changes.
[0081] Hybrid Non-Technical Loss Detection System
[0082] Embodiments of the present invention provide a method for detecting fraudulent SM based on consumption data of a smart grid. To detect fraudulent activity, embodiments of the present invention use ARIMA forecasting. ARIMA is a statistical model that predicts future trends using time series data (non-patent literature
[13] ). The ARIMA method is an improvement on the moving-average approach. Moving-average forecasting is performed using the previous time series data point that has the strongest correlation with the current value.
[0083] The predicted value is predicted through the analysis of the autocorrelation function (ACF) and partial autocorrelation function (PACF) synthesized with the order input for the ARIMA function (non-patent literature
[29] ). The ARIMA model has the condition that the time series data must be fixed. In fixed time series data, the statistical mean, variance, and autocorrelation are all constant over time.
[0085] ARIMA-based NTL detection system
[0086] The following 'HNTLD Scheme Supporting Observer Smart Meters (HNTLDS-ObSM)' describes a hybrid NTL scheme with ObSM for detecting energy fraud, which uses ARIMA to predict readings based on available ObSM readings. If ObSM is not part of the infrastructure, the hybrid NTL detection scheme without ObSM support (HNTLDS-WObSM), described below, can be used for energy fraud detection, allowing ARIMA to learn from data that is strongly correlated with previous similar meter readings. ARIMA can be used in both methods.
[0087] In the HNTLDS-ObSM method, ARIMA is used as an interpolation method to predict missing reads that match global ObSM reads. The HNTLDS-WObSM method predicts missing SM reads by considering acceptable limits.
[0088] FIG. 1 is a flowchart illustrating an NTL fraud detection method according to an embodiment of the present invention. The entire sequence of steps of an ARIMA-based NTL detection system used to predict NTL fraud cases is illustrated in FIG. 1.
[0089] Referring to FIG. 1, a method according to an embodiment of the present invention begins by checking the availability of readings from observer meters or observer smart meters, which are considered as the sum of a set of monitored smart meters (S100). Then, after extracting readings using an ARIMA time series forecasting function with a related time series forecasting function, a random forest machine learning model can terminate the final step of detecting fraud.
[0090] Through these steps, surge patterns can be detected in the usage readings of a smart meter, and the accuracy of detecting fraud in normal readings can be improved compared to conventional NTL detection methods. In addition, because the method according to the embodiment of the present invention is dynamic, utility companies can increase or decrease the fraudulent activity threshold.
[0091] Below, an NTL fraud detection method is described in detail with reference to Fig. 1.
[0092] First, check if an observer meter reading is available (S100). An observer meter is a device that measures the total energy of a power distribution line and is used to compare with the sum of smart meters. If an observer meter reading is available (Yes), calculate the difference between the observer meter reading and the smart meter reading (S102). By comparing the difference in data between the observer meter (upper meter) and the smart meter (lower meter), estimate the amount of loss (e.g., missing power consumption). Afterward, check if it is non-fraudulent (S104). If the result of the data comparison is reasonable (Yes), determine that it is non-fraudulent (normal user) and terminate.
[0093] If it is confirmed that it is not fraudulent (No), the individual data row where fraud occurred is identified (S106). The point in time at which an anomaly occurred in the smart meter suspected of fraudulent activity is identified. Subsequently, the smart meter dataset is separated into non-fraudulent and suspected (fraudulent) (S108). The dataset is separated into normal users and suspected users. Subsequently, the approximate reading of the suspected (fraudulent) smart meter is predicted using ARIMA (S110). The power consumption that should have been recorded if it were normal is predicted using the ARIMA model.
[0094] Subsequently, fraudulent smart meters are identified by comparing the ARIMA predicted value with the actual reported reading (S112). If the actual reported value is significantly lower than the predicted value, it is highly likely that it was intentionally manipulated. Afterward, the fraudulent smart meter is reported (S114) and the process terminates. Smart meters suspected of being fraudulent are reported to the utility company or monitoring system.
[0095] If observer meter readings are unavailable in S100 (No), an approximate reading of a fraudulent smart meter is predicted using ARIMA (S110). Then, the fraudulent smart meter is identified by comparing the ARIMA predicted value with the actual reported reading (S112). Then, the fraudulent smart meter is reported (S114) and the process terminates.
[0097] Hybrid NTL Detect Scheme Supporting ObSM (HNTLDS-ObSM)
[0098] The system according to an embodiment of the present invention assumes that there is one ObSM for each set of n SMs. The total electricity supplied to these SMs is recorded in the ObSM. The mathematical model variables and symbols are as follows.
[0099] T j : Automatic SM data extraction from the j-th period
[0100] E j : Total ObSM electrical readings in the j-th period
[0101] e i,j : SM i Regarding period T j Electrical data recorded on ObSM
[0102] r i,j : Period T j At SM i Electricity reading consumed at
[0103] R j : Reading r measured at each i-th (1, 2, 3, ..., n) SM i,j vectors of
[0104] c i,j : Period T j At SM i A constant of a function representing
[0105] C: An n-dimensional vector consisting of n coefficients for n smart meters
[0107] Time-frequency T j Total reading E of ObSM j is the electricity e supplied to all SMs connected to it. i,j It is equal to the sum of. Ideally, each SM (r i,j If the reading of ) is predicted through approximation, ObSM (e i,j It is identical to the reading of the electricity supplied through ). The individual reading r of SM in the j-th collected measurement j,j (Here, i=1, 2, 3, ..., n) is available, but in ObSM, the reading e of each SM i,j cannot be used.
[0108] SM i related to (r i,j , e i,j For the ) pair, the function is y i = f i (r)(here, f i (r i,j ) = e i,j It can be expressed as ). SM having m samples i ARIMA approximation e for i,j To generate, the approximate ARIMA fitting function passes through all m points. This can be expressed as Equation 1 below. Here, c k,i is a constant, f i (r i,j ) = e i,j (1≤j≤m), r k is SM i These are individual SM readings for .
[0109]
[0110] c k,i The value is e k,i / r k,i It is calculated as. All SM readings can be represented through the R(m,n) matrix.
[0111] Mathematical Equation 2 is the j-th ObSM reading E in the case of no fraud or adjusted SM. j It represents. This value row and It is the sum of the values multiplied by the columns.
[0112]
[0113] E in mathematical formula 3 T and R -1 If you input, you can calculate a vector C of size n. Here, n represents the number of SMs.
[0114]
[0115] If all C vector values are close to 1, all n SMs are considered non-fraudulent. On the other hand, if some items of the C vector are far from 1, there exists at least one fraudulent SM.
[0116] Fraudulent rows are identified by calculating mathematical formula 4 for each i-th row.
[0117]
[0118] (Here, (which is the tolerance and is typically 0.02), if the SM reading for the corresponding row is non-fraudulent data. On the other hand, If so, the SM reading for row i is fraud data. As such, the SM data is divided into two parts. That is, the first is the Before_fraud_data from row 1 to row i-1. i (Data prior to the fraudulent activity), and the other is Suspected_fraudulent_data from row i to the last row. iIt is (data on suspected fraudulent activity).
[0119] Before_fraud_data i An ARIMA model is trained using [this], and predicted values are calculated using the generated model. A fraud threshold is constructed using Chebyshev's Theorem (non-patent literature
[30] ). The trained model ARIMA-FIT i Using, an ARIMA-forecast consisting of (n-1) ARIMA prediction values for the smart meter SM i Creates a set. This is SM's Suspected_fraudulent_data i The expected value e corresponding to i,j (i.e., the actual measured value of the section suspected of fraud). Fraudulent activity is identified using mathematical formula 5.
[0120]
[0121] If the result of mathematical formula 5 consistently yields True for a period longer than the given period, SM is normal and not fraudulent. Otherwise, SM is fraudulent.
[0123] Hybrid NTL detection scheme that does not support ObSM (HNTLDS-WObSM)
[0124] This technique works when ObSM data is unavailable. ARIMA time series analysis is used to identify fraudulent activity. First, a trained model is constructed based on customers' known non-fraudulent consumption behaviors using an ARIMA approach with consumption values from previous periods.
[0125] The consumption value for the current period is predicted using an ARIMA prediction model. It is assumed that there is a significant difference between the ARIMA predicted reading and the reported SM reading for the current period. In this case, it is presumed to be NTL fraud and reported using Chebyshev's theorem. The ARIMA prediction method is identical to the HNTLDS-ObSM method. An ARIMA prediction step is applied to generate an expected ARIMA predicted reading based on previously learned reading behavior, and to determine whether it is NTL fraud.
[0126] To compare the ARIMA-based system according to an embodiment of the present invention with a conventional system, an ARIMA fraud detection function is applied to the SGCC dataset (non-patent document
[31] ). For each SM reading, an ARIMA prediction value (SM reading, ARIMA prediction, trend, seasonality, residual) is generated. This feature set is passed to a Random Forest (RF) machine learning supervised approach (non-patent document
[14] ,
[32] ) to generate a classification model that classifies the SM reading as fraudulent or normal.
[0127] Random forests are efficient machine learning algorithms that merge multiple decision trees for more optimal classification results. Since random forests operate and run efficiently on large datasets, they are suitable for learning fraudulent behavior on the SGCC dataset with provided labels. Random forests are not biased by outliers (non-patent literature
[33] ). In addition, they can be easily scaled in parallel by performing explicit feature selection. Random forests have a low tendency for overfitting and are effective in balancing imbalanced datasets (non-patent literature
[34] ). A random forest-based HNTLDS model according to an embodiment of the present invention is illustrated in FIG. 2.
[0129] <Performance Evaluation>
[0130] Experiment setup
[0131] The system according to an embodiment of the present invention was implemented using RStudio version 1.2.5033 (non-patent document
[35] ). Computation tests were performed on a gaming workstation equipped with an 11GB NVIDIA graphics card and a Windows 10 OS. In an embodiment of the present invention, two publicly available datasets were used for energy fraud identification and performance comparison as follows.
[0132] - SGCC dataset: This dataset is from the State Grid Corporation of China (non-patent literature
[31] ). This data concerns daily SM electricity readings for 42,372 consumers for 1,035 days starting from January 1, 2014. Of the 42,372 consumers, 3,615 were marked as fraudulent use. The columns of the SGCC dataset are consumer number, fraud flag (1 for fraudulent record, 0 for non-fraud), reporting time (15-minute intervals), and reading.
[0133] - UCI dataset: This dataset is a public dataset related to individual household power consumption data in the Sceaux, Paris region, and can be downloaded from non-patent literature
[36] ,
[37] . It contains 2,075,259 records from December 2006 to December 2010. The attributes of this dataset are SM number, usage value, and reporting time (1-minute intervals). It also includes ObSM reporting data.
[0135] performance metric
[0136] Various performance indicators (non-patent literature
[32] ) were considered to compare the performance of the system according to the embodiment of the present invention with that of conventional systems (non-patent literature [5],
[16] –
[20] ). Accuracy (ACC), precision, and false positive rate (FPR) are frequently used measurements to assess the performance of energy fraud detection systems. Area under the curve (AUC) measures how perfectly fraud cases are classified from normal cases.
[0137] The proportion of cases correctly detected as fraud is called recall or sensitivity (Se). This measures the model's ability to accurately detect fraud cases. The proportion of cases correctly identified as non-fraud is called specificity (Sp), which measures the model's ability to correctly detect non-fraud cases. The proportion of correctly predicted fraud cases that are actually identified as fraud is called precision (Pr). Furthermore, the balance between sensitivity and specificity is evaluated by calculating the F1 score. A higher F1 score indicates that the model identifies as many fraud cases as possible, resulting in a lower FPR.
[0139] Performance Analysis
[0141] HNTLDS-ObSM results
[0142] Two experiments were conducted on sample data from the UCI dataset. In the first experiment, the data was accurate, and no SMs committed fraud. In the second experiment, fraudulent behavior was observed among some SMs. This can be confirmed by comparing the sum of SM data over a specific time sample period with the ObSM data.
[0143] Results of Experiment #1: The samples considered in this experiment correspond to the samples analyzed in non-patent literature
[36] and consist of 20 samples, each consisting of 10 SM readings and corresponding ObSM readings. The SM readings and their corresponding ObSM readings are within the range of ±0.02. When the method according to the embodiment of the present invention is executed, the results for the C vector items are (0.99, 0.99, 1.00, 0.99, 0.99, 1.00, 1.00, 1.00, 1.00, 1.00). And, since all indicators of the C vector are nearly equal to 1 for each SM, the corresponding SM is considered normal and no fraudulent activity occurred.
[0144] Results of Experiment #2: NTL Detection in Fraud Data
[0145] The samples considered in this experiment are taken from the UCI dataset and include 10 SM readings and 30 records of readings responded to by ObSM. Some SM readings are fraudulent (Non-patent literature
[36] ). The result of the C vector in Equation 3 is (1.53, 2.27, -0.25, 2.85, 0.95, 0.82, 0.85, 1.70, 0.78, 0.74), and this value is far from 1. Therefore, the difference between the sum of each row and the corresponding ObSM measurement is calculated, and the following vector is generated. (0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1)
[0146] This indicates that fraud began at the 25th item. An ARIMA model is fitted for the first 24 rows of each SM, and ARIMA forecasting is performed on the remaining 6 items using the generated model. Applying Equation 4, SM3 and SM6 are identified as fraud because their readings do not fall within the 3σ range.
[0147] Fraudulent and non-fraudulent SMs are illustrated in Fig. 3. SM3 and SM6 are identified as fraudulent SMs, and their operation is illustrated in Fig. 3. The graphs of SM3 and SM6 show values that continuously fall outside the acceptable range from the 25th to the 30th reading (below μ-σ), followed by a sharp rise in readings. However, the values of the other graphs are located between μ-σ and μ+σ.
[0149] HNTLDS-WObSM results
[0150] Experiment #3: NTL Detection Using a Pure ARIMA Model
[0151] This experiment demonstrates that using ARIMA allows for the prediction of behavior in a specific period based on usage data from previous periods. The difference between ARIMA predictions and observed usage must not exceed a specific limit, namely ±0.02.
[0152] Figure 4 is a diagram illustrating the detection of fraudulent and non-fraudulent behavior in SM using ARIMA. Figure 4a is a diagram visually showing the learned model approach, illustrating fraudulent behavior in SM. The reported values and predicted values do not follow the same pattern. The red line (reported values) moves downward from the lowest limit (green line). Figure 4b illustrates non-fraudulent behavior in SM. The predicted values and reported values are located within the same range.
[0154] Performance comparison between conventional methods and HNTLDS-WObSM
[0155] The generated dataset is used to train and test ARIMA results using a machine learning random forest model (non-patent literature [7]). Essentially, a random forest model is a combination of various decision trees. The random model can provide superior detection performance while preventing overfitting compared to existing decision trees.
[0156] The hyperparameters of the Random Forest are selected via grid search. The data is divided into training, testing, and validation sets in proportions of 70%, 15%, and 15%, respectively. The model is constructed with default training parameters (training_frame = train, stopping_rounds = 5, stopping_tolerance = 0.001, stopping_metric = AUC, seed = 1, balance_classes = FALSE, nfolds = 10), the number of epochs is set to 20, the random seed is set to 1, and the number of trees is set to 250. The value of n, the number of estimators, is 256, and the maximum depth is 8.
[0157] The time required to train the model on a 1GB NVIDIA graphics card gaming workstation is 25 minutes. The comparison focused on six metrics: sensitivity (recall), accuracy, specificity, AUC, F1 score, and precision. RMSE is used to calculate the mean variance between the actual (label) and the pre-predicted value for each row of data. RMSE measures the model's prediction error and indicates how well the model's predictions match the data of the actual labels. Comparing the achieved Random Forest classification performance results with human judgment, a model accuracy of 98% was achieved.
[0158] Figure 5 is a diagram showing a comparison between a system according to an embodiment of the present invention (Proposed method-Random Forest) and a conventional system (CNN-GRU-PSO, Linear SVM, CNN, RUSBoost, WADCNN, DERUSBOOST) considering only the accuracy performance indicator.
[0159] Table 1 and Figure 6 show a comparison between a system according to an embodiment of the present invention and a conventional system, taking into account all six performance indicators.
[0160]
[0161] The proposed method using ARIMA with Random Forest in Table 1 and the proposed method in Fig. 6 represent embodiments of the present invention. Refer to non-patent literature
[18] for CNN-GRU-PSO, non-patent literature [5] for Linear SVM, non-patent literature
[19] for CNN, non-patent literature
[16] for adasyn, non-patent literature [5] for RUSBoost, non-patent literature
[20] for WADCNN, and non-patent literature
[17] for DERUSBOOST.
[0162] As shown in Fig. 6, the model according to an embodiment of the present invention exhibits better performance than conventional models and can accurately identify fraud cases with trivial RMSE errors. The most notable sensitivity is 98.2%. High specificity (specificity 99.3%) for recognizing non-fraud cases with trivial errors, a very high F1 score (0.984), and an accuracy of 0.98 indicate a good level for detecting fraud. Additionally, the reported Matthew Correlation Coefficient (MCC) is 0.97, and the final RMSE is 0.02.
[0163] In Table 1, the DERUSBOOST classifier optimized with the differential evolution meta-heuristic technique is robust against outliers and can handle high-dimensional data, but it may overfit to the data and be computationally expensive in high-dimensional data. The specificity performance (99.6%) is nearly identical, but the method according to the embodiment of the present invention achieves 98%, surpassing it.
[0164] In non-patent literature
[19] , the authors used Paillier cryptography with CNNs to detect abnormal consumption SM data behaviors using stochastic gradient descent (SGD) optimization tools. The accuracy of this method was only 93%. While CNNs are adept at finding features and are computationally efficient, the model can overfit if not normalized. This paper reports only the accuracy, which indicates the model's ability to reliably classify all results.
[0165] Since FP (false positive) and FN (false negative) must be considered, better metrics must be used to quantify model performance. The accuracy of the WADCNN scheme is only 86.3%, making it the least effective. While it can handle high-dimensional data and is resilient to outliers, it can overfit if not properly managed.
[0166] A hybrid method utilizing CNN, PSO, and GRUCNN-GRU-PSO (non-patent literature
[18] ) combines a CNN with a gated recurrent neural network optimized by Particle Swarm Optimization. This method is robust against outliers and can handle high-dimensional data. Although it is computationally expensive, this can be overcome through normalization. The performance of this model is reported as 89%, 87.3%, and 87% in terms of AUC, accuracy, and F1 score, respectively. This is insufficient to describe fraud cases because sensitivity has not been reported. Therefore, the model's ability to detect fraud cases is ambiguous.
[0167] Adasyn is an effective method for handling class balance, enabling greater learning in fraudulent situations where it is difficult to learn without bias. In non-patent literature [5], the authors used linear SVM, non-linear SVM, and a multi-layer sensory neural network to detect NTL in unbalanced data. This technique has excellent specificity but low sensitivity (54.2%) to non-fraudulent cases. Linear SVM can efficiently handle high-dimensional data. However, if not normalized, the problem of overfitting may occur.
[0168] FIG. 7 is a graph showing a performance comparison between an embodiment of the present invention and a conventional method. FIG. 7a is a figure showing the FPR (false positive rate) AUC (ROC-AUC) relative to the TPR (true positive rate), and FIG. 7b is a figure showing the recall (sensitivity) AUC (PR-AUC) relative to accuracy. In FIG. 7, The Proposed ARIMA-RF represents an embodiment of the present invention.
[0169] As illustrated in FIG. 7b, the method according to an embodiment of the present invention surpasses the conventional method in all performance indicators. This indicates that the model according to an embodiment of the present invention accurately classified fraudulent electricity consumers and legitimate electricity consumers. FIG. 7b compares the method according to an embodiment of the present invention with the conventional method using PR-AUC and ROC-AUC.
[0170] Deep learning algorithms such as Adasyn (non-patent document
[16] ), CNN-GRU-PSO (non-patent document
[18] ), WADCNN (non-patent document
[20] ), and DERUSBOOST (non-patent document
[17] ) have superior performance compared to previous machine learning algorithms such as linear SVM (non-patent document [5]) based on AUC values (0.95, 0.89, 0.76, 0.89, and 0.80, respectively). However, the method according to the embodiment of the present invention has a superior performance compared to these, with an AUC value of 0.978. The Adasyn method is powerful for classifying high-dimensional information and automatically extracts features from the input, so it ranked second highest in AUC ranking following the method according to the embodiment of the present invention.
[0171] FIG. 8 is a block diagram showing a loss detection system of a smart grid according to an embodiment of the present invention.
[0172] Referring to FIG. 8, the memory (110) is connected to one or more processors (130) and can store codes that cause the processors (130) to control a loss detection system of a smart grid when executed by the processors (130). The memory (110) may include magnetic storage media or flash storage media, but the scope of the invention is not limited thereto. The memory (110) may include internal memory and / or external memory and may include volatile memory such as DRAM, SRAM, or SDRAM, non-volatile memory such as OTPROM (one time programmable ROM), PROM, EPROM, EEPROM, mask ROM, flash ROM, NAND flash memory, or NOR flash memory, non-transient computer-readable storage media such as SSD, CF (compact flash) card, or SD card.
[0173] In addition, various information necessary within the scope of achieving the purpose of the present disclosure may be stored in the memory (110), and the information stored in the memory (110) may be updated as it is received from a server or external device or input by a user.
[0174] The communication unit (120) may provide a communication interface necessary to provide transmission and reception signals between external devices (including servers) in the form of packet data in conjunction with a network. Additionally, the communication unit (120) may be a device including hardware and software necessary to transmit and receive signals, such as control signals or data signals, through wired or wireless connections with other network devices.
[0175] The processor (130) can receive various data or information from an external device connected through the communication unit (120) and can also transmit various data or information to the external device. In addition, the communication unit (120) may include at least one of a WiFi module, a Bluetooth module, a wireless communication module, and an NFC module.
[0176] The input unit (140) is an input interface for collecting various data applied to the smart grid loss detection system. The data can be input by a user or obtained from a server. Additionally, the input unit (140) may receive user commands to control the operation of the smart grid loss detection system and may include, for example, a microphone, a touch display, etc.
[0177] The output unit (150) is an output interface to which the results of the loss detection system of the smart grid are output, and may include, for example, a display.
[0178] The processor (130) can control the overall operation of the smart grid loss detection system. Specifically, the processor (130) is connected to the configuration of the smart grid loss detection system including the memory (110) as described above, and can control the overall operation of the smart grid loss detection system by executing at least one command stored in the memory (110) as described above.
[0179] The processor (130) can be implemented in various ways. For example, the processor (130) can be implemented as at least one of an Application Specific Integrated Circuit (ASIC), an embedded processor, a microprocessor, hardware control logic, a hardware finite state machine (FSM), or a digital signal processor (DSP).
[0180] A loss detection system for a smart grid according to an embodiment of the present invention may include a memory (110) and a processor (130) connected to the memory (110) and configured to execute computer-readable commands included in the memory (110).
[0181] The processor (130) performs the following operations: (A) checking whether the reading of the observer meter (ObSM) is available; (B) if the reading of the observer meter is available, (B-1) calculating the difference between the reading collected from a plurality of smart meters (SM) and the reading of the observer meter; (B-2) determining whether the smart meter is a fraudulent smart meter based on the difference between the reading of the smart meter and the reading of the observer meter, wherein if the difference between the reading of the smart meter and the reading of the observer meter is less than a predetermined threshold, it is determined to be normal, and if it exceeds the threshold, the individual row in which fraud has occurred is identified; (B-3) after the individual row is identified, separating the smart meter dataset into normal data and suspicious data; (B-4) predicting the approximate reading of the suspicious smart meter using ARIMA (auto-regressive integrated moving average); and (B-5) identifying the fraudulent smart meter by comparing the predicted value of ARIMA with the actual reported reading.
[0182] Meanwhile, (C) if the reading of the observer meter is not available, (C-1) an operation to predict an approximate reading of the suspected smart meter using ARIMA (auto-regressive integrated moving average); and (C-2) an operation to identify a fraudulent smart meter by comparing the predicted value of ARIMA with the actual reported reading.
[0183] A detailed description of the operation performed by the processor (130) is omitted as it overlaps with the description above.
[0184] As described above, the present invention has been explained by specific details such as specific components, limited embodiments, and drawings; however, these are provided merely to aid in a more comprehensive understanding of the invention, and the invention is not limited to the above embodiments. A person skilled in the art to which the invention pertains will be able to make various modifications and variations within the scope of the essential characteristics of the invention. Accordingly, the concept of the present invention should not be limited to the described embodiments, and all technical concepts that are equivalent to or have equivalent variations to the claims set forth below, as well as the claims themselves, should be interpreted as being included within the scope of the rights of the present invention. Furthermore, each of the above embodiments may be combined and operated as needed. Explanation of the symbols
[0185] 110: Memory 120: Communications Department 130: Processor 140: Input section 150: Output section
Claims
Claim 1 (A) a step of checking whether the readings of an observer meter (ObSM) are available; (B) if the readings of the observer meter are available, (B-1) a step of calculating the difference between the readings collected from a plurality of smart meters (SM) and the readings of the observer meter; (B-2) a step of determining whether the smart meter is a fraudulent smart meter based on the difference between the readings of the smart meter and the readings of the observer meter, wherein if the difference between the readings of the smart meter and the readings of the observer meter is less than a predetermined threshold, it is determined to be normal, and if it exceeds the threshold, an individual row in which fraud has occurred is identified; (B-3) after the individual rows are identified, a step of separating the dataset of the smart meter into normal data and suspicious data; (B-4) a step of predicting the approximate readings of the suspicious smart meter using ARIMA (auto-regressive integrated moving average); A method for detecting loss in a smart grid, comprising: (B-5) a step of identifying a fraudulent smart meter by comparing the predicted value of the ARIMA with the actual reported reading; and (C) if the reading of the observer meter is unavailable, (C-1) a step of predicting an approximate reading of the suspected smart meter using ARIMA (auto-regressive integrated moving average); and (C-2) a step of identifying a fraudulent smart meter by comparing the predicted value of the ARIMA with the actual reported reading. Claim 2 A method for detecting loss in a smart grid according to claim 1, wherein the loss is a non-technical loss. Claim 3 In claim 1, the step of identifying the individual row in step (B-2) above is performed by calculating the following mathematical formula 4 for each row, and go ( is the allowable deviation, and if it is 0.02 or less, it is a normal reading. go A smart grid loss detection method in which a reading less than is a fraudulent reading.[Equation 4] (here, r j,j period T j At SM i It is the electricity reading consumed at, and E j is the total ObSM electrical reading in the j-th period.) Claim 4 A method for detecting loss in a smart grid according to claim 1, wherein the threshold value is constructed using Chebyshev's theorem. Claim 5 A method for detecting loss in a smart grid according to claim 1, wherein the step of identifying the fraudulent smart meter in steps (B-5) and (C-2) is performed by a random forest machine learning model. Claim 6 In claim 1, the step of identifying the fraudulent smart meter in steps (B-5) and (C-2) is a smart grid loss detection method performed by the following mathematical formula 5, wherein if the result of the mathematical formula 5 continuously yields true for a predetermined period or longer, the smart meter is normal. [Mathematical Formula 5] (Here, ARIMA_forecast i is the value predicted by the ARIMA model for smart meter i, and Suspected_fraudulent_data i is the actual measurement value for the section suspected of fraud regarding smart meter i.) Claim 7 The apparatus comprises: a memory; and a processor connected to the memory and configured to execute computer-readable commands contained in the memory, wherein the processor comprises: (A) an operation to check whether a reading of an observer meter (ObSM) is available; (B) if the reading of the observer meter is available, (B-1) an operation to calculate the difference between the reading collected from a plurality of smart meters (SM) and the reading of the observer meter; (B-2) an operation to determine whether the smart meter is a fraudulent smart meter based on the difference between the reading of the smart meter and the reading of the observer meter, wherein if the difference between the reading of the smart meter and the reading of the observer meter is less than a predetermined threshold, it is determined to be normal, and if it exceeds the threshold, an operation to identify an individual row in which fraud has occurred; (B-3) after the individual row is identified, an operation to separate the dataset of the smart meter into normal data and suspicious data; (B-4) an operation to predict an approximate reading of the suspicious smart meter using ARIMA (auto-regressive integrated moving average); A smart grid loss detection system capable of executing the following: (B-5) identifying a fraudulent smart meter by comparing the predicted value of the ARIMA with the actual reported reading; (C) if the reading of the observer meter is unavailable, (C-1) predicting an approximate reading of the suspected smart meter using ARIMA (auto-regressive integrated moving average); and (C-2) identifying a fraudulent smart meter by comparing the predicted value of the ARIMA with the actual reported reading. Claim 8 In paragraph 7, the above loss is a non-technical loss, a loss detection system for a smart grid. Claim 9 In claim 7, the operation of identifying the individual row in the above (B-2) operation is performed by calculating the following mathematical formula 4 for each row, and go ( is the allowable deviation, and if it is 0.02 or less, it is a normal reading. go A smart grid loss detection system in which a reading of less than is a fraudulent reading.[Equation 4] (here, r j,j period T j At SM i It is the electricity reading consumed at, and E j is the total ObSM electrical reading in the j-th period.) Claim 10 In claim 7, the above threshold is configured using Chebyshev's theorem, a loss detection system for a smart grid. Claim 11 In claim 7, the operation of identifying the fraudulent smart meter in the above (B-5) and above (C-2) operations is performed by a random forest machine learning model, in a smart grid loss detection system. Claim 12 In claim 7, the operation of identifying the fraudulent smart meter in the above (B-5) and above (C-2) is a smart grid loss detection method performed by the following mathematical formula 5, wherein if the result of the above mathematical formula 5 continuously yields true for a predetermined period or longer, the smart meter is normal, a smart grid loss detection system.[Mathematical Formula 5] (Here, ARIMA_forecast i is the value predicted by the ARIMA model for smart meter i, and Suspected_fraudulent_data i is the actual measurement value for the section suspected of fraud regarding smart meter i.) Claim 13 A program stored on a computer-readable storage medium for executing a loss detection method performed by a loss detection system of a smart grid, comprising: (A) an operation to check whether a reading value of an observer meter (ObSM) is available; (B) if the reading value of the observer meter is available, (B-1) an operation to calculate the difference between the reading value collected from a plurality of smart meters (SM) and the reading value of the observer meter; (B-2) an operation to determine whether the smart meter is a fraudulent smart meter based on the difference between the reading value of the smart meter and the reading value of the observer meter, wherein if the difference between the reading value of the smart meter and the reading value of the observer meter is less than a predetermined threshold, it is determined to be normal, and if it exceeds the threshold, an operation to identify an individual row in which fraud has occurred; (B-3) after the individual row is identified, an operation to separate the dataset of the smart meter into normal data and suspicious data; (B-4) an operation to predict an approximate reading value of the suspicious smart meter using ARIMA (auto-regressive integrated moving average); A program stored on a computer-readable storage medium, comprising: (B-5) an operation to identify a fraudulent smart meter by comparing the predicted value of the ARIMA with the actual reported reading; and (C) if the reading of the observer meter is unavailable, (C-1) an operation to predict an approximate reading of the suspected smart meter using ARIMA (auto-regressive integrated moving average); and (C-2) an operation to identify a fraudulent smart meter by comparing the predicted value of the ARIMA with the actual reported reading. Claim 14 In paragraph 13, the above loss is a program stored on a computer-readable storage medium, which is a non-technical loss. Claim 15 In paragraph 13, the operation of identifying the individual row in the above (B-2) operation is performed by calculating the following mathematical formula 4 for each row, and go ( is the allowable deviation, and if it is 0.02 or less, it is a normal reading. go A program stored on a computer-readable storage medium, which is a fraudulent reading if less than [Equation 4]. (here, r j,j period T j At SM i It is the electricity reading consumed at, and E j is the total ObSM electrical reading in the j-th period.) Claim 16 In paragraph 13, the above threshold is a program stored on a computer-readable storage medium, constructed using Chebyshev's theorem. Claim 17 In paragraph 13, the operation of identifying the fraudulent smart meter of the above (B-5) and above (C-2) is a program stored on a computer-readable storage medium, which is performed by a random forest machine learning model. Claim 18 In paragraph 13, the operation of identifying the fraudulent smart meter in the above (B-5) and above (C-2) is a smart grid loss detection method performed by the following Equation 5, wherein the smart meter is normal if the result of Equation 5 continuously yields true for a predetermined period or longer, and is a program stored on a computer-readable storage medium.[Equation 5] (Here, ARIMA_forecast i is the value predicted by the ARIMA model for smart meter i, and Suspected_fraudulent_data i is the actual measurement value for the section suspected of fraud regarding smart meter i.)