A Scaling Attack Detection Method Considering the Diversity of Electricity Consumption Patterns in the Advanced Metering Infrastructure of Smart Grid
By analyzing the characteristics of scaling attacks in the smart grid, using the decision tree and Kmeans method to extract the power consumption interval and perform discrete processing, the problem that smart meters in the smart grid are vulnerable to data integrity attacks is solved, effectively detecting scaling attacks, and improving data security.
Patent Information
- Application Number
- CN202211105584.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-05
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-09-05
AI Technical Summary
Smart meters in smart grids are susceptible to data integrity attacks, and existing detection methods are difficult to effectively detect a single attack type, especially when users have multiple power consumption modes.
By analyzing the characteristics of various data integrity attack types, selecting the most representative scaling attacks, and using machine learning methods such as decision trees and Kmeans, extracting the power consumption interval and performing discrete processing to detect scaling attacks.
This method can effectively detect scaling attacks when the user has multiple power modes, improve detection performance and ensure data security.
Smart Images

Figure CN115641227B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of smart grid, and mainly relates to a scaling attack detection method considering the diversity of electricity consumption patterns in the advanced metering infrastructure of the smart grid, providing a basis for ensuring its data security. Background Art
[0002] The smart grid makes up for the deficiencies of the traditional power grid, such as unidirectional information flow, single energy use, and low user participation, by integrating the advanced metering infrastructure (AMI) and various distributed energy sources, realizing the modernization of the power system. As an important component of AMI, smart meters play an important role in the process of information transmission. It collects and uploads the electricity consumption information of user households and receives the decision-making information after the power company makes electricity consumption decisions, enabling users to supply and use electricity more reasonably and economically. However, since smart meters are deployed in an open network environment, the risk of being attacked by network attacks is increased, making it a potential target for data integrity attacks. Attackers launch data integrity attacks on smart meters to tamper with user electricity consumption information, aiming to profit for themselves and harm the safe operation of the smart grid.
[0003] At present, the methods for detecting data integrity attacks in the smart grid are mainly divided into three categories: state-based, game theory-based, and classification-based methods. State-based detection methods require a large amount of additional costs for power companies, such as software and hardware costs, increased operation / training costs, etc.; in game theory-based methods, how to formulate a utility function is a challenging task, and ultimately all attacks cannot be detected. With the continuous maturity of machine learning technology, most attack behaviors can be detected at a moderate cost by using appropriate classifiers and datasets. Therefore, classification-based detection research has gradually become the mainstream. In existing classification schemes, when detecting AMI attacks, there is a problem of uniformly identifying and detecting all attack types. Since different attack types have different characteristics, there is currently no algorithm that can effectively detect all attack types. When the scenario is a single attack type, existing methods cannot show good detection performance. Therefore, how to design a detection method for a single data integrity attack type needs further research.
[0004] In summary, we believe that the key to solving the above problems is how to select representative attack types and design appropriate algorithms according to their characteristics. This patent analyzes the characteristics of various data integrity attack types, selects the most representative scaling attack for detection, extracts the electricity consumption interval based on machine learning methods such as decision trees and Kmeans, discretizes it using the interval, and uses the electricity consumption magnitude as the basis for feature selection, and then proposes a scaling attack detection method considering the diversity of electricity consumption patterns in AMI. Summary of the Invention
[0005] To overcome the problems existing in the above technologies, the present invention proposes a scaling attack detection method considering the diversity of electricity consumption patterns in AMI.
[0006] A scaling attack detection method for an advanced metering infrastructure in a smart grid considering the diversity of electricity consumption patterns according to the present invention includes the following steps:
[0007] (1) Represent the user's electricity consumption data as c = [c1, c2…c j …c sum T , where sum represents the total number of days of data collection, and c j = [c j-1 , c j-2 …c j-h …c j-24 represents the electricity consumption data on the j-th day, and c j-h represents the electricity consumption in the h-th time period on the j-th day; check for missing values in the electricity consumption data. When the number of missing values in a piece of electricity consumption data is less than 6 and there are no consecutive missing values, the average electricity consumption of the previous and the next time periods is used to fill in the value, as shown in Equation (1);
[0008]
[0009] If there are consecutive missing values, the average value of the entire piece of data is used to fill in the missing values; when the number of missing values in a piece of electricity consumption data
[0010] is greater than 6, this piece of data is represented as unavailable;
[0011] (2) Using the Kmeans algorithm with the Euclidean distance as the distance metric, divide all electricity consumption data c = [c1, c2…c j …c sum into K electricity consumption patterns C = [C 1 , C 2 …C k …C K , where C k represents the k-th electricity consumption pattern set, as shown in Equation (2);
[0012]
[0013] where, c k d_h represents the electricity consumption in the h-th time period on the d-th day in the k-th electricity consumption pattern;
[0014] (3) Extract the k electricity consumption pattern intervals I = [I1, I2…I k …I K in step 2, Ik The interval representing the k-th power consumption mode set, as shown in Equation (3);
[0015] I k =[min k , max k (3)
[0016] where min k represents the minimum power consumption per unit time period in the k-th power consumption mode, that is, for any c k d_h , there is min k ≤c k d_h ; max k represents the maximum power consumption per unit time period in the k-th power consumption mode, that is, for any c k d_h , there is max k ≥c k d_h ;
[0017] (4) According to the scaling attack model c j-h * =γ h c j-h , γ h =random(0.1, 0.8) to generate the attack data c j * , where γ h represents a random number between 0.1 and 0.8 that changes with time, c j-h * represents the original data value c j-h multiplied by γ h to obtain the attack data value, c j * represents the data under attack; mix the attack data with the normal data, perform binarization using the power consumption interval, and discretize the values of the power consumption data within the interval I to 0 and those outside the interval I to 1, as shown in Equation (4);
[0018]
[0019] (5) Let the number of features discretized to 0 in a piece of data be Z; all the feature values of the normal data c j belong to the interval I, so Z of the normal data = 24; while for the attack data c j * , more than half of the feature values fall outside the interval I after the attack, so Z * of the attack data < 12. Let represent the set of 24 time period features, T inand T out represent two different subsets; when the time period is h, if the number of attack data with the time period eigenvalue in the interval I is not less than the number of attack data outside the interval I, that is, when it satisfies |c j-h * ∈I|≥|c j * |-|c j-h * ∈I|, then T h ∈T in , where it represents the number of data; if the number of attack data with the time period eigenvalue in the interval I is less than the number of attack data outside the interval I, that is, when it satisfies , then T h ∈T out ;
[0020] (6) Calculate the empirical conditional entropy of each feature and select the feature; the goal of the decision tree is to continuously find the time period with the largest information gain, that is, to find g(D, T h ) max that satisfies g(D, T h ) max > g(D, T s ), s≠h; where g(D, T h ) = H(D) - H(D|T h ), represents the information gain of the feature T h to the data set D, H(D) represents the empirical entropy of the data set D, and H(D|T h ) represents the empirical conditional entropy of the feature T h to the data set D; the empirical entropy of the data set D is a fixed value, so only the feature corresponding to the minimum empirical conditional entropy needs to be found, that is, H(D|T h ) min that satisfies H(D|T h ) min < H(D|T s ), s≠h;
[0021] For T h ∈T in , as the power consumption data continuously increases:
[0022] |D i=1 |→0, |D i=0 |→|D (5)
[0023] D i represents the set of power consumption data with the value of i (0 or 1) at the time period T h ; when the data set is balanced:
[0024]
[0025] D i,l represents the set of electricity consumption data with a value of i and an electricity consumption data category of l (0 or 1) at time period T, and calculates T h ∈ T h ∈ T in in
[0026] the empirical conditional entropy of the feature, as shown in Equation (7);
[0027]
[0028] (7) Calculate T h ∈ T out the empirical conditional entropy of the feature in T, for T h ∈ T out , as the electricity consumption data continues to increase:
[0029]
[0030] When the data set is balanced:
[0031]
[0032] Calculate T h ∈ T out the empirical conditional entropy of the feature in T, as shown in Equation (10);
[0033]
[0034] (8) According to Equation (7) and Equation (10), the decision tree uses the features in T h ∈ T out to construct a tree for detection; when the newly collected data is detected, it is checked in turn whether the discrete value corresponding to the time period in the data has a value of 1 according to the selected features in the tree. If it exists, the data is detected as attack data; if it does not exist, it is detected as normal data.
[0035] The method of the present invention can ensure that when the user has multiple electricity consumption patterns, scaling attacks can be effectively detected. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flowchart of a scaling attack detection method considering the diversity of electricity consumption patterns in AMI;
[0037] Figure 2 is a performance graph of a scaling attack detection method considering the diversity of electricity consumption patterns in AMI;
[0038] Figure 3 is a comparison graph of the F1 scores of the scaling attack detection method considering the diversity of electricity consumption patterns in AMI and the scaling attack detection method without considering the diversity of electricity consumption patterns;
[0039] Figure 4 It is a comparison chart of the F1 scores of the scaling attack detection method considering the diversity of electricity consumption patterns in AMI with the KNN method and the Naive Bayes method. Specific implementation manner
[0040] The present invention will be further described in detail below with reference to the accompanying drawings:
[0041] A scaling attack detection method for an advanced metering infrastructure in a smart grid considering the diversity of electricity consumption patterns according to this embodiment includes the following steps:
[0042] (1) Represent the user's electricity consumption data as c = [c1, c2... c j …c sum T , where sum represents the total number of days of data collection, and c j = [c j-1 , c j-2 …c j-h …c j-24 represents the electricity consumption data on the j-th day, and c j-h represents the electricity consumption in the h-th time period on the j-th day; check for missing values in the electricity consumption data. When the number of missing values in a piece of electricity consumption data is less than 6 and there are no consecutive missing values, the average electricity consumption of the previous and the next time periods is used to fill in the value, as shown in Equation (1);
[0043]
[0044] If there are consecutive missing values, the average value of the entire piece of data is used to fill in the missing values; when the number of missing values in a piece of electricity consumption data is greater than 6, the piece of data is represented as unavailable;
[0045] (2) Using the Kmeans algorithm with the Euclidean distance as the distance metric, divide all the electricity consumption data c = [c1, c2... c j …c sum into K electricity consumption patterns C = [C 1 , C 2 …C k …C K , where C k represents the k-th electricity consumption pattern set, as shown in Equation (2);
[0046]
[0047] Among them, c k d_h represents the electricity consumption in the h-th time period on the d-th day in the k-th electricity consumption pattern;
[0048] (3) Extract the k power consumption pattern intervals I = [I1, I2… I k … I K , where I k represents the interval of the k-th power consumption pattern set, as shown in Equation (3);
[0049] I k = [min k , max k (3)
[0050] Among them, min k represents the minimum power consumption per unit time period in the k-th power consumption pattern, that is, for any c k d_h , there is min k ≤ c k d_h ; max k represents the maximum power consumption per unit time period in the k-th power consumption pattern, that is, for any c k d_h , there is max k ≥ c k d_h ;
[0051] (4) According to the scaling attack model c j-h * = γ h c j-h , γ h = random(0.1, 0.8) to generate the attack data c j * , where γ h represents a random number between 0.1 and 0.8 that changes with time, c j-h * represents the original data value c j-h multiplied by γ h to obtain the attack data value, c j * represents the data to be attacked; mix the attack data with the normal data, binarize using the power consumption interval, and discretize the values of the power consumption data within the interval I to 0, and those outside the interval I to 1, as shown in Equation (4);
[0052]
[0053] (5) Let the number of features discretized to 0 in a piece of data be Z; all feature values of the normal data c j belong to the interval I, so Z = 24 for the normal data; while for the attack data c j *More than half of the eigenvalues fall outside the interval I after an attack, so the Z of the attack data * <12. Let represent the set of 24 time period features, T in and T out represent two different subsets; when the time period is h, if the number of attack data with eigenvalues in the interval I is not less than the number of attack data outside the interval I, that is, when |c j-h * ∈I|≥|c j * |-|c j-h * ∈I|, T h ∈T in , where represents the number of data; if the number of attack data with eigenvalues in the interval I is less than the number of attack data outside the interval I, that is, when is satisfied, T h ∈T out ;
[0054] (6) Calculate the empirical conditional entropy of each feature and select the feature; the goal of the decision tree is to continuously find the time period with the largest information gain, that is, to find g(D, T h ) max that satisfies g(D, T h ) max >g(D, T s ), s≠h; where g(D, T h ) = H(D) - H(D|T h ), represents the information gain of the feature T h to the dataset D, H(D) represents the empirical entropy of the dataset D, and H(D|T h ) represents the empirical conditional entropy of the feature T h to the dataset D; the empirical entropy of the dataset D is a constant value, so only the feature corresponding to the minimum empirical conditional entropy needs to be found, that is, H(D|T h ) min that satisfies H(D|T h ) min <H(D|T s ), s≠h;
[0055] For T h ∈T in , as the power consumption data continues to increase:
[0056] |D i=1 |→0, |D i=0 |→|D| (5)
[0057] D iRepresents the set of electricity consumption data with value i (0 or 1) at time period T h ; When the data set is balanced:
[0058]
[0059] D i,l Represents the set of electricity consumption data with value i and electricity consumption data category l (0 or 1) at time period T h , calculate the empirical conditional entropy of the features in T h ∈T in , as shown in Equation (7);
[0060]
[0061] (7) Calculate the empirical conditional entropy of the features in T h ∈T out ; For T h ∈T out , as the electricity consumption data continues to increase:
[0062]
[0063] When the data set is balanced:
[0064]
[0065] Calculate the empirical conditional entropy of the features in T h ∈T out , as shown in Equation (10);
[0066]
[0067] (8) According to Equation (7) and Equation (10), the decision tree uses the features in T h ∈T out to construct a tree for detection; when the newly collected data is detected, check whether there is a value of 1 in the discrete value corresponding to the time period in the data in turn according to the selected features in the tree. If there is, the data is detected as attack data; if not, it is detected as normal data.
[0068] Suppose a user has 3 electricity consumption patterns, and each pattern has 1586 pieces of data. Each piece of data records the electricity consumption of the user every hour of the 24 hours of a day. After checking and processing the missing values, the three patterns are distinguished by Kmeans clustering, and some data are randomly selected from each pattern to generate attack data. Set the ratio of the training set to the test set to 7:3, then 854 pieces of data are randomly selected from each pattern to generate attack data. Finally, the number of the training set is 5124, and the number of the test set is 2196.
[0069] Based on the above settings, we conducted two sets of simulations using the False Positive Rate (FPR), False Negative Rate, and F1-score as evaluation metrics:
[0070] False Positive Rate: The proportion of normal data detected as attack data among all normal data.
[0071] False Negative Rate: The proportion of attack data detected as normal data among all attack data.
[0072] F1-score: The weighted average of model precision and detection rate, taking both precision and detection rate into account.
[0073] Figure 2 The model performance when the attack intensity in the test set increases from 10% to 80% is presented. It can be seen that when the attack proportion in the test set increases from 10% to 80%, the F1-score is [95.7% 96.23% 96.32% 96.4% 96.38% 96.41% 96.46% 96.52%], the FPR is [0.18% 0.20% 0.18% 0.20% 0.20% 0.19% 0.19% 0.20%], and the FNR is [6.69% 6.51% 6.69% 6.63% 6.78% 6.79% 6.75% 6.67%]. Therefore, regardless of the attack intensity, the method has a high F1-score, a low FPR, and a low FNR, which verifies the effectiveness of the method. Figure 3 A comparison graph of the F1-score between the method of the present invention and the scaling attack detection method that does not consider the diversity of electricity consumption patterns is presented. It can be seen that when not considering the diversity of electricity consumption patterns, if the minimum value is taken as the discrete threshold, the F1-score is much lower than that of the method of the present invention; when the average value is taken as the threshold, although the F1-score increases as the number of attack data in the test set increases, it never exceeds the method of the present invention. Figure 4 A comparison graph of the F1-score between the method of the present invention and the KNN method and the Naive Bayes method is presented. It can be seen that the F1-score of the KNN method is slightly higher than 92%, the F1-score of the Naive Bayes method is slightly higher than 94%, while the F1-score of the method of the present invention is always higher than 96%, which verifies the high efficiency of the method of the present invention.
Claims
1. A scaling attack detection method considering the diversity of electricity consumption patterns in the advanced metering infrastructure of smart grid, characterized in that Including the following steps: (1) Represent the user's electricity consumption data as c = [c1, c2... c j ... c sum T , where sum represents the total number of days for data collection, and c j = [c j-1 , c j-2 ... c j-h ... c j-24 represents the electricity consumption data for the j-th day, and c j-h represents the electricity consumption for the h-th time period on the j-th day; check for missing values in the electricity consumption data. When the number of missing values in a piece of electricity consumption data is less than 6 and there are no consecutive missing values, use the average of the electricity consumption in the previous and next time periods of this value for filling, as shown in Equation (1); If there are consecutive missing values, the average value of the entire piece of data is used to fill in the missing values; when the number of missing values in a piece of electricity consumption data is greater than 6, this piece of data is represented as unavailable; (2) Using the Kmeans algorithm with the Euclidean distance as the distance metric, all electricity consumption data c = [c1, c2... c j ... c sum is divided into K electricity consumption patterns C = [C 1 , C 2 ... C k ... C K , where C k represents the k-th electricity consumption pattern set, as shown in Equation (2); Among them, c k d_h represents the electricity consumption in the h-th time period of the d-th day in the k-th electricity consumption mode; (3) Extract the k power consumption pattern intervals I = [I1, I2... I k ... I K , where I k represents the interval of the k-th power consumption pattern set, as shown in Equation (3); I k = [min k , max k (3) where, min k represents the minimum power consumption per unit time period in the k-th power consumption mode, that is, for any c k d_h , there is min k ≤c k d_h ; max k represents the maximum power consumption per unit time period in the k-th power consumption mode, that is, for any c k d_h , there is max k ≥c k d_h ; (4) According to the scaling attack model c j-h * = γ h c j-h , γ h = random(0.1, 0.8) to generate the attack data c j * , where γ h represents a random number ranging from 0.1 to 0.8 that changes with time, and c j-h * represents the original data value c j-h multiplied by γ h to obtain the attack data value, and c j * represents the data under attack; mix the attack data with the normal data, use the power consumption interval for binarization, discretize the values of the power consumption data within the interval I to 0, and discretize the values outside the interval I to 1, as shown in Equation (4); (5) Let the number of features with discrete value 0 in a piece of data be Z; for normal data c j all feature values belong to the interval I, so for normal data, Z = 24; while for attack data c j * more than half of the feature values fall outside the interval I after an attack, so for attack data, Z * < 12. Let represent the set of 24 time - period features, T in and T out represent two different subsets; when the time period is h, if the number of attack data with feature values within the interval I is not less than the number of attack data with feature values outside the interval I, that is, when it satisfies |c j-h * ∈I|≥|c j * |-|c j-h * ∈I|, then T h ∈T in , where || represents the number of data; if the number of attack data with feature values within the interval I is less than the number of attack data with feature values outside the interval I, that is, when it satisfies , then T h ∈T out ; (6) Calculate the empirical conditional entropy of each feature and select the feature; the goal of the decision tree is to continuously find the time period with the largest information gain, that is, to find g(D, T h ) max satisfying g(D, T h ) max > g(D, T s ), s ≠ h; where g(D, T h ) = H(D) - H(D|T h ), representing the information gain of feature T h on the data set D, H(D) represents the empirical entropy of the data set D, and H(D|T h ) represents the empirical conditional entropy of feature T h on the data set D; the empirical entropy of the data set D is a constant value, so only the feature corresponding to the minimum empirical conditional entropy needs to be found, that is, H(D|T h ) min satisfying H(D|T h ) min < H(D|T s ), s ≠ h; For T h ∈ T in , with the continuous increase of power consumption data: |D i=1 |→0,|D i=0 |→|D| (5) D i represents the set of electricity consumption data with a value of i (0 or 1) at time period T h ; when the data set is balanced: D i,l represents the set of power consumption data with a value of i and a power consumption data category of l (0 or 1) at time period T h , and calculate the empirical conditional entropy of the features in T h ∈T in as shown in Equation (7); (7) Calculate T h ∈ T out Calculate the empirical conditional entropy of the features in T. For T h ∈ T out , as the electricity consumption data continues to increase: When the dataset is balanced: Calculate T h ∈ T out The empirical conditional entropy of the features in it is shown in Equation (10); (8) According to Equation (7) and Equation (10), the decision tree uses the features in T h ∈T out to construct a tree for detection; when the newly collected data is detected, check in sequence whether the discrete value in the corresponding time period of the data has a value of 1 according to the selected features in the tree. If it exists, the data is detected as attack data; if it does not exist, it is detected as normal data.
Citation Information
Patent Citations
Bayesian-based enterprise risk classification model construction method
CN110543904A
Small sample driven abnormal power consumption data set construction method and module
CN113190595A