Data integrity attack defense method for forwarding node based on trust in advanced measurement system
By using the isolated forest algorithm iForest in the smart grid for abnormal detection, and combining the power consumption period characteristics and the change mode of the meter reading behavior to calculate the trust value, the problem of insufficient accuracy in the calculation of trust value in the existing technology is solved, and more efficient data transmission path security optimization is achieved.
Patent Information
- Application Number
- CN202510536875.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-01
AI Technical Summary
In the data integrity attack defense against smart meters as forwarding nodes, it is difficult to effectively use meter energy data for abnormal detection, and the characteristics of electricity use periods and the change patterns of meter reading behaviors are not fully considered, resulting in insufficient accuracy of the calculation of trust value.
IForest is used to detect abnormalities by using the isolated forest algorithm iForest, and combined with the multi-dimensional characteristics of the energy data of electricity meter, an isolated forest is constructed to identify abnormal data. At the same time, the trust value weight is dynamically adjusted according to the characteristics of the power consumption period, and the change mode of the electricity meter is analyzed through the sliding window method to calculate the overall trust value. Ultimately, by weighted summing, the comprehensive trust value is obtained, and path selection is optimized to reduce the risk of data attacks.
It improves the accuracy of abnormal detection, enhances the accuracy of trust value calculation, optimizes the security of data transmission paths, and effectively reduces the risk of data being attacked on forwarding nodes.
Smart Images

Figure CN120238366A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of Smart Grid, and particularly relates to a trust-based data integrity attack defense method for forwarding nodes in an Advanced Metering Infrastructure. Based on considering the time characteristics of electricity consumption and the change patterns of electricity meter reading behaviors, data integrity protection is carried out by ensuring the security of electricity meter energy data transmission. Background Art
[0002] As a new generation of power infrastructure, the Smart Grid realizes the cross-border integration of theory and technology by integrating advanced information and communication technologies with traditional power systems. With the support of the Advanced Metering Infrastructure (AMI) of a high-speed two-way communication network, by means of innovative sensing and metering means, modern equipment, intelligent control algorithms, and scientific decision-making support systems, the performance of the power grid in terms of system reliability, security, economic benefits, operation efficiency, and environmental sustainability has been comprehensively improved. As an important part of the Smart Grid, AMI consists of Smart Meters (SMs), Data Concentrators (DCs), Utilities, and the high-speed two-way communication network between them. AMI realizes the two-way interaction of communication and power data between the source node (Smart Meter) and the Utility, including the data uplink stage and the data downlink stage, enabling the Utility to monitor, meter, and control the electricity consumption of users. The data uplink stage includes data measurement and transmission by the Smart Meter - data forwarding by the Smart Meter and the Data Concentrator - data reception by the Utility: As the core component of AMI, each Smart Meter periodically measures user data, including electricity meter communication data such as packet loss rate, message sending and receiving frequencies, etc.; electricity meter forwarding data such as packet forwarding behaviors, forwarding frequencies, forwarding paths, transmission delays, etc.; electricity meter energy data including temperature, time, "whether someone is at home", electricity meter readings (users' electricity consumption), etc. User data is measured and sent by the source node (Smart Meter), forwarded by the forwarding nodes (Smart Meter and Data Concentrator), and received by the Utility. The data downlink stage includes data sending by the Utility - data forwarding by the Smart Meter and the Data Concentrator - data reception by the Smart Meter, that is, the Utility forwards data to the source node via the forwarding nodes through a path opposite to the uplink stage. The data collected by the Utility through the data uplink stage is widely used in intelligent decision-making such as real-time electricity prices and power energy scheduling, and is the basis and driving factor for the Utility to make decisions, playing a decisive role in decision-making.
[0003] However, smart meters are exposed in an open environment that is easily accessible to users, and communicate with other smart grid components through an open two-way wireless network. This openness provides new attack points and attack paths for attackers. Attackers can easily capture and violate smart meters through data integrity attacks, and tamper with user data during the measurement, forwarding process, resulting in electricity price, power supply, and power consumption decisions made by power companies deviating from normal values, causing serious consequences such as supply-demand imbalance, user power outage, and damage to user economic interests. In the entire data routing process of user data measurement and sending-forwarding-receiving, when the smart meter acts as a forwarding node, it is the bridge for two-way interaction between the source node and the power company. Moreover, compared with the source node and the power company, the number of forwarding nodes is larger. Therefore, deploying a secure user data integrity attack defense model for forwarding nodes is the top priority for ensuring the safe operation of the power grid and has become a research hotspot.
[0004] Existing user data integrity attack defense models for forwarding nodes are mainly divided into the following types: the method based on dynamic grouping, the method based on redundant verification, and the method based on trust evaluation. Although the method based on dynamic grouping is highly effective in data integrity protection, its implementation process is complex. Especially in the process of key update and group management, the computational overhead of the system is large. The method based on redundant verification enhances the fault tolerance of the model by introducing additional redundant nodes and error retransmission mechanisms. However, this method requires additional hardware investment and may bring communication and computational burdens. In contrast, the method based on trust evaluation dynamically adjusts the trust value by evaluating the behavior of forwarding nodes, with relatively small computational and communication overheads and avoiding excessive hardware dependence, having the lightweight characteristic, and thus gradually becoming the mainstream solution. Although existing methods can select secure forwarding smart meters based on trust values and play a role in defending against user data integrity attacks on forwarding smart meters, there are still the following two problems:
[0005] (1) In terms of anomaly detection:
[0006] First of all, existing methods for data anomaly detection are mainly divided into those based on node communication data (such as packet loss rate, message sending and receiving frequency, etc.) and those based on node forwarding data (such as packet forwarding behavior, forwarding frequency, forwarding path, transmission delay, etc.). Only a small number of methods introduce meter energy data for detection. However, compared with node communication data and forwarding data, meter energy data has unique advantages in revealing user behavior and electricity consumption patterns, so it has higher application value in anomaly detection.
[0007] In addition, most existing methods adopt methods based on node behavior anomaly detection (such as adaptive trust attribute acquisition methods, trust attribute anomaly judgment methods based on Mahalanobis distance) and methods based on distributed methods (such as consensus-based distributed methods, distributed collaborative detection-based methods, etc.). However, anomaly detection methods such as Mahalanobis distance and majority voting algorithm perform well in dealing with a small amount of low-dimensional data, but the detection accuracy is not high when facing large-scale high-dimensional data such as electricity meter energy data.
[0008] (2) In terms of trust evaluation:
[0009] First of all, a small number of existing methods involve anomaly detection of electricity meter energy data, but usually do not fully consider the influence of the time period characteristics of electricity consumption on trust evaluation. The electricity consumption situation shows peak electricity consumption periods, normal electricity consumption periods, and low electricity consumption periods within a day. The electricity meter reading characteristics in these periods are different, and the corresponding trust value calculation should be adjusted accordingly. If the time period characteristics are ignored in the trust value calculation, it may be difficult to accurately identify or process the reading fluctuations caused by electricity consumption changes. For example, during peak electricity consumption periods, the electricity meter readings may naturally increase, while during low consumption periods, they will decrease. If these normal reading fluctuations are not included in the trust value calculation logic, it is easy to cause misjudgment of the readings, thus affecting the accuracy of trust value calculation.
[0010] In addition, the calculation methods for evaluating trust values in existing methods are mainly divided into: direct trust value calculation, where a node calculates the trust value based on the direct interaction behavior of other nodes, such as Bayesian methods, window-based schemes, etc.; overall trust value calculation, where a node calculates the trust value of neighbor nodes based on the recommendations or feedback of other nodes, such as Dempster-Shafer theory, etc.; and comprehensive trust value calculation, where a node comprehensively considers direct trust, overall trust, and other factors, such as historical trust, penalty factors, etc., to calculate the comprehensive trust value. However, these methods do not comprehensively and fully consider all possible change patterns of electricity meter reading behaviors when calculating trust values, that is, there are "persistently abnormal smart meters", "gradually recovering smart meters", and "highly stable smart meters", making it impossible to accurately reflect and process the dynamic changes of electricity meter reading behaviors. For example, for a "persistently abnormal smart meter", if its possible recovery is not considered, its trust value may be overly reduced, thus affecting the accuracy of trust value calculation.
[0011] Therefore, it is necessary to propose a trust-based data integrity attack defense method for forwarding nodes that makes full use of electricity meter energy data, has high anomaly detection accuracy, and considers the time period characteristics of electricity consumption and the change patterns of electricity meter reading behaviors.
[0012] In summary, how to make full use of the energy data of the electricity meter and select an algorithm with high anomaly detection accuracy, while considering the characteristics of the electricity consumption period and the change pattern of the electricity meter reading behavior, is the key to solving the above problems. Summary of the Invention
[0013] To overcome the problems existing in the above technologies, the present invention proposes a trust-based data integrity attack defense method for forwarding nodes in an advanced metering infrastructure.
[0014] To achieve the above object, the technical solution adopted by the present invention is:
[0015] A trust-based data integrity attack defense method for forwarding nodes in an advanced metering infrastructure, comprising the following steps:
[0016] (1) Data collection:
[0017] Each user i, i = 1, 2,..., m has an intelligent electricity meter SM i , i = 1, 2,..., m, where m is the total number of users; each intelligent electricity meter SM i regularly collects user data Message in units of period t i_t , including communication data Traffic i_t , forwarding data Forward i_t and energy data Energy i_t , Energy i_t includes hour Hour i_t , minute Minute i_t , temperature Temperature i_t , "whether someone is at home" Occupier i_t and electricity meter reading Data i_t ; in addition, each intelligent electricity meter SM i can also obtain the electricity meter energy data of each of its neighbor intelligent electricity meters SM j , j = 1, 2,..., n i , where n i represents the number of neighbor intelligent electricity meters of the intelligent electricity meter SM i ; the neighbor intelligent electricity meters of SM i refer to other intelligent electricity meters that are one-hop reachable from the intelligent electricity meter SM i in the same network topology; the n i neighbor intelligent electricity meters SM i , j = 1, 2,..., n j of each intelligent electricity meter SM i constitute its neighbor node set N i ;
[0018] (2) Anomaly detection:
[0019] (2-1) Form a data set:
[0020] Each smart meter SM i Obtains the energy data of each of its neighbor nodes SM from the power company j , j = 1, 2,..., n i The energy data for the previous d periods, denoted as Energy j_c , c = t - d, t - d + 1,..., t - 1, to form the data set D j , j = 1, 2,..., n i ;
[0021] (2-2) Generate isolation trees:
[0022] Each smart meter SM i For each neighbor node SM j , j = 1, 2,..., n i Of the electricity meter energy data set D j Perform sub-sampling, SM i Randomly select ψ j , ψ j ∈ D j Entries of data to obtain the sub-data set D j '; Then, randomly select a feature set Q j = {Hour j , Minute j_t , Temperature j_t , Occupier j_t , Data j_t} in a feature q j_t ; Then randomly select a value s j Within the value range of this feature as the splitting point, and use the data with feature values less than this value as the left child, otherwise as the right child; Repeat this process until each generated isolation tree reaches the specified depth h j Or the number of samples contained in each leaf node does not exceed the threshold n j_max ; j_min ;
[0023] (2-3) Construct an isolation forest:
[0024] Repeat the process of splitting sub-trees described in (2-2), and gradually construct iTree j ~ iTree j_1 ~ iTree j_jnumIsolated trees, where jnum represents the number of iTrees; finally, all the constructed isolated trees are merged to form a complete isolated forest iForest j ;
[0025] (2-4) Calculate the anomaly score:
[0026] Calculate the average path length of each data Energy in the isolated forest j_c as shown in Equation (1):
[0027]
[0028] where H j (Energy j_c ) is the average path length of data Energy j_t in the j-th tree, and E(H(Energy j_c ) is the average path length of data Energy j_c ; then, normalize the path to obtain the normalized average path length C(d) as shown in Equation (2):
[0029]
[0030] where d is the number of data in the electricity meter energy dataset D j and H d is the harmonic number providing the normalization benchmark; finally, substitute into Equation (3) to obtain the anomaly score of data Energy j_c :
[0031]
[0032] where S(Energy j_c ,d) is the anomaly score of the electricity meter energy data Energy j_c ,Energy j_c ∈D j and its value range is [0-1];
[0033] (2-5) Judge the abnormal data:
[0034] According to the anomaly score S(Energy j_c ,d), by setting the threshold θ j ,j = 1,2,...n i judge whether the data Energy j_c is abnormal data: when S(Energy j_c ,d)≥θ j , the data Energy j_c is judged as abnormal data, when S(Energyj_c , d) < θ j When the data Energy j_c is judged as normal data;
[0035] (3) Trust evaluation:
[0036] (3 - 1) Direct trust value calculation:
[0037] First, considering the differences in electricity consumption in different time periods of each day, a day is divided into three time periods: peak electricity consumption period peak, normal electricity consumption period off_peak, and low electricity consumption period low. Secondly, count the number of abnormal data and the total number of abnormal data in each time period of the data in D j respectively, which are denoted as N peak , N off_peak , N low , N total , and then calculate the probabilities P1, P2, and P3 of abnormal data occurring in the peak electricity consumption period, normal electricity consumption period, and low electricity consumption period according to formula (4):
[0038]
[0039] Thirdly, based on P1, P2, and P3 obtained from formula (4), perform weighted summation to obtain the direct trust value DT i , as shown in formula (5):
[0040] DT i = P1×α + P2×β + P3×γ, (5)
[0041] where α, β, and γ are the weights of the three time periods;
[0042] Specifically, α, β, and γ are calculated by formula (6) from the initial weights α1, β2, γ3:
[0043]
[0044] where weight_func is the weight adjustment function, as shown in formula (7):
[0045] weight_func(p) = 1 - P; (7)
[0046] (3 - 2) Overall trust value calculation:
[0047] First, obtain all d data in the D j dataset; then, set a sliding window Window, and set the size of the sliding window Window length to w, and the step size Window λ is also set to w, then there are a total of A window is used, and the number l of abnormal data in each window is counted; Window length = Window λ The window size and step size are matched with the time characteristics of the data, so as to balance the computational complexity and the accuracy of anomaly detection; Secondly, a threshold μ is set. For each window Window j , j = 1, 2, 3..., o. When the number l of abnormal data in the window is greater than the threshold μ, it is marked as the change mode Model1: "Intelligent meter with continuous anomaly", indicating that the meter has been in an untrustworthy state for a long time; If the number l of abnormal data is greater than 0 but not greater than μ, it is marked as the change mode Model2: "Intelligent meter gradually recovering", indicating that the meter is recovering to the normal state; If the number l of abnormal data is 0, it is marked as the change mode Model3: "Intelligent meter with high stability", indicating that the meter has always performed well; Thirdly, different weights W i are assigned to the three change modes, which are respectively W i1 , W i2 and W i3 , satisfying:
[0048]
[0049] Then, the overall trust value of each meter is calculated using Equation (9):
[0050]
[0051] Specifically, the weight Win i of each window is multiplied by its exponential decay function value and accumulated to obtain the overall trust value of the window; Finally, the average value of the overall trust values of all windows is taken as the final overall trust value; Among them, the exponential decay formula is the penalty factor, l i is the number of abnormal data in each window, and different θ i are assigned to the three change modes, denoted as θ i1 , θ i2 and θ i3 , satisfying:
[0052] θ i1 > θ i2 > θ i3 > 0; (10)
[0053] (3-3) Comprehensive trust value calculation:
[0054] The calculation of the comprehensive trust value is a weighted sum of the direct trust value and the overall trust value, so as to obtain the final trust value of each intelligent meter; The calculation of the comprehensive trust value is as follows:
[0055] TV i = λ × DT i + (1 - λ) × OT i , (11)
[0056] Where λ is a parameter between 0 and 1, used to balance the influence of the direct trust value and the overall trust value on the comprehensive trust value;
[0057] (4) Optimal secure path selection:
[0058] (4-1) Each smart meter SM i constructs the optimal secure path p starting from itself as the start path node i ; Let i = 1, and the start path node be x1 = SM i , that is, the optimal secure path is p i = x1;
[0059] (4-2) Determine whether there is a data concentrator DC in the neighbor set N i of the current node; if it exists, use DC as the next-hop path node and add it to the optimal secure path p i , that is, p i = p i ∪{DC}, and the path selection ends;
[0060] (4-3) If there is no data concentrator in the neighbor set, traverse the nodes in the set N i , where the comprehensive trust value of each neighbor node SM j is denoted as TV j , select the node with the maximum comprehensive trust value TV j that is not in the current path p i as the next-hop path node x i+1 , and add it to the path p i , that is, p i = p i ∪{x i+1}, x i+1 is denoted as:
[0061]
[0062] Then, filter out the nodes that already exist in the path from the neighbor node set, that is, N i = N i / {x i+1}, update i = i + 1, and return to (4-2).
[0063] First, in the data collection stage, multi-dimensional electricity meter energy data is collected, including data such as temperature, time, "whether someone is at home", and electricity meter readings. The Isolation Forest algorithm iForest is used to perform anomaly detection and identification based on these electricity meter energy data. Then, in the trust evaluation stage, a day is divided into peak electricity consumption periods, normal electricity consumption periods, and low electricity consumption periods. The weights are dynamically adjusted according to the occurrence probability of abnormal data in each time period. On this basis, the direct trust value is calculated by means of weighted summation. Secondly, the sliding window method is introduced to dynamically analyze three change patterns of the electricity meter, namely, "smart electricity meters with continuous anomalies", "smart electricity meters gradually recovering", and "smart electricity meters with high stability". Weights are assigned based on the characteristics of different change patterns, and the exponential decay function is combined as a penalty factor to adjust the influence degree of each change pattern on the trust value. By accumulating the product of the weight and the penalty factor within the sliding window and taking the average, the overall trust value is finally calculated. Finally, the direct trust value and the overall trust value are weighted and summed to obtain the comprehensive trust value, comprehensively improving the accuracy of trust value calculation. Finally, in the path selection stage, smart electricity meters with higher comprehensive trust values are preferentially selected as forwarding nodes to construct the optimal secure path, thereby reducing the risk of data integrity attacks on user data in the forwarding nodes and maximizing the security of the data transmission process.
[0064] Therefore, the present invention has the following advantages:
[0065] In terms of anomaly detection, the present invention introduces electricity meter energy data for detection and uses the Isolation Forest algorithm iForest to improve the detection accuracy under large-scale high-dimensional data; in terms of trust evaluation, the time period characteristics of electricity consumption and the change patterns of electricity meter reading behaviors are fully considered, thereby improving the accuracy of trust value calculation, further optimizing the performance of the optimal secure path with the highest calculated path security value, and finally realizing the secure transmission of data. Description of the Drawings
[0066] Figure 1 is the flowchart of the trust-based data integrity attack defense method for forwarding nodes in AMI;
[0067] Figure 2 is the optimal secure path diagram of the 88th group;
[0068] Figure 3 is the optimal secure path diagram of the 104th group;
[0069] Figure 4 is the optimal secure path diagram of the 122nd group;
[0070] Figure 5 is the optimal secure path diagram of the 245th group;
[0071] Figure 6 is the optimal security path graph of the 367th group;
[0072] Figure 7 is the optimal security path graph of the 421st group;
[0073] Figure 8 is the comparison graph of the trust - based data integrity attack defense method for forwarding nodes and the non - trust - based random path selection method in AMI;
[0074] Figure 9 is the performance graph of the trust - based data integrity attack defense method for forwarding nodes in AMI under different anomaly ratios;
[0075] Figure 10 is the comparison graph of FNR of the isolation forest algorithms iForest, LOF, and SVM;
[0076] Figure 11 is the comparison graph of FPR of the isolation forest algorithms iForest, LOF, and SVM;
[0077] Figure 12 is the comparison graph of F1 of the isolation forest algorithms iForest, LOF, and SVM. Detailed implementation manner
[0078] The present invention will be further described in detail below with reference to the accompanying drawings:
[0079] As Figure 1 shown, the trust - based data integrity attack defense method for forwarding nodes in AMI described in this embodiment has the following specific steps:
[0080] (1) Data collection:
[0081] Each user i, i = 1, 2,..., m has a smart meter SM i , i = 1, 2,..., m, where m is the total number of users; each smart meter SM i regularly collects user data Message i_t at a period of t, including communication data Traffic i_t , forwarding data Forward i_t and energy data Energy i_t , Energy i_t includes hour Hour i_t , minute Minute i_t , temperature Temperature i_t , "whether someone is at home" Occupier i_t and meter reading Data i_t; In addition, each smart meter SM i can also obtain the electricity meter energy data of each of its neighbor smart meters SM j , j = 1, 2,..., n i from the power company, where n i represents the number of neighbor smart meters of the smart meter SM i ; The neighbor smart meters of SM i refer to other smart meters that are one-hop reachable from the smart meter SM i in the same network topology; Each smart meter SM i 's n i neighbor smart meters SM j , j = 1, 2,..., n i constitute its neighbor node set N i ;
[0082] (2) Anomaly detection:
[0083] (2 - 1) Form a data set:
[0084] Each smart meter SM i obtains the energy data of each of its neighbor nodes SM j , j = 1, 2,..., n i for the previous d periods, denoted as Energy j_c , c = t - d, t - d + 1,..., t - 1, to form a data set D j , j = 1, 2,..., n i ;
[0085] (2 - 2) Generate isolation trees:
[0086] Each smart meter SM i performs sub - sampling on the electricity meter energy data set D j of each neighbor node SM i , j = 1, 2,..., n j ; SM i randomly selects ψ j , ψ j ∈ D j data items to obtain a sub - data set D j '; Then, randomly select a feature set Q j = {Hour j , Minute j_t , Temperature j_t , Occupier j_t , Data j_t} in the sub - data set D j_t and a feature q j; Then randomly select a value s within the value range of this feature j As the splitting point, the data with feature values less than this value are used as the left child, otherwise as the right child; repeat this process until each generated isolated tree reaches the specified depth h j_max Or the number of samples contained in each leaf node does not exceed the threshold n j_min ;
[0087] (2-3) Construct the isolation forest:
[0088] Repeat the process of splitting the subtree described in (2-2), and gradually construct iTree j ~iTree j_1 Isolated trees, where jnum represents the number of iTrees; finally, merge all the constructed isolated trees to form a complete isolation forest iForest j_jnum ; j ;
[0089] (2-4) Calculate the anomaly score:
[0090] Calculate the average path length of each data Energy j_c in the isolation forest, as shown in Equation (1):
[0091]
[0092] Among them, H j (Energy j_c ) is the average path length of the data Energy j_t in the j-th tree, and E(H(Energy j_c ) is the average path length of the data Energy j_c ; then, standardize the path to obtain the standardized average path length C(d), as shown in Equation (2):
[0093]
[0094] Among them, d is the number of data in the electricity meter energy dataset D j , and H d is the harmonic number providing the standardization benchmark; finally, substitute it into Equation (3) to obtain the anomaly score of the data Energy j_c :
[0095]
[0096] Among them, S(Energy j_c ,d) is the electricity meter energy data Energy j_c ,Energy j_c ∈Dj The abnormal score, whose value range is [0 - 1];
[0097] (2 - 5) Judging abnormal data:
[0098] According to the abnormal score S(Energy j_c , d), by setting the threshold θ j , j = 1, 2,... n i Judge whether the data Energy j_c is abnormal data: When S(Energy j_c , d) ≥ θ j , the data Energy j_c is judged as abnormal data. When S(Energy j_c , d) < θ j , the data Energy j_c is judged as normal data;
[0099] (3) Trust evaluation:
[0100] (3 - 1) Calculating the direct trust value:
[0101] First, considering the differences in electricity consumption in different time periods of each day, a day is divided into three time periods: peak electricity consumption period peak, off - peak electricity consumption period off_peak, and low - valley electricity consumption period low. Second, count the number of abnormal data and the total number of abnormal data in each time period for the data in D j , which are respectively denoted as N peak , N off_peak , N low , N total , and then calculate the probabilities P1, P2, and P3 of abnormal data occurring in the peak electricity consumption period, off - peak electricity consumption period, and low - valley electricity consumption period according to formula (4):
[0102]
[0103] Third, based on P1, P2, and P3 obtained from formula (4), perform weighted summation to obtain the direct trust value DT i , as shown in formula (5):
[0104] DT i = P1×α + P2×β + P3×γ, (5)
[0105] Among them, α, β, and γ are the weights of three time periods. Since the number of abnormal data in a time period reflects the occurrence of abnormal events (i.e., data is judged as abnormal data) in different time periods, different weights need to be assigned to different time periods. In addition, the weight values of α, β, and γ need to be dynamically adjusted according to different probabilities, so that the time period with a large abnormal probability obtains a small weight value;
[0106] Specifically, α, β, and γ are calculated by the initial weights α1, β2, and γ3 through Equation (6):
[0107]
[0108] Among them, weight_func is the weight adjustment function, as shown in Equation (7):
[0109] weight_func(p) = 1 - P; (7)
[0110] (3 - 2) Calculation of the overall trust value:
[0111] First, obtain all d data in the D j dataset; then, set a sliding window Window, and the size of the sliding window Window length is set to w, and the step size Window λ is also set to w, then there are windows, and count the number of abnormal data l in each window; Window length = Window λ makes the window size and step size match the time characteristics of the data, and can achieve a balance between computational complexity and abnormal detection accuracy; second, set a threshold μ. For each window Window j , j = 1, 2, 3..., o, when the number of abnormal data l in the window is greater than the threshold μ, it is marked as the change pattern Model1: "Continuously abnormal smart meter", indicating that the meter has been in an untrusted state for a long time; if the number of abnormal data l is greater than 0 but not greater than μ, it is marked as the change pattern Model2: "Gradually recovering smart meter", indicating that the meter is recovering to normal; if the number of abnormal data l is 0, it is marked as the change pattern Model3: "High-stability smart meter", indicating that the meter has always performed well; third, assign different weights W i to the three change patterns, which are W i1 、W i2 and W i3 , satisfying:
[0112]
[0113] Among them, W i2The maximum value indicates that the influence degree of change pattern Model2 is the largest, because it is necessary to promptly reflect the recovery situation of the electricity meter, thus avoiding too low trust value due to short-term anomalies; W i1 The minimum value indicates that the influence degree of change pattern Model1 is relatively small, because it is necessary to impose a greater penalty on continuous abnormal behaviors to reflect their negative impact on the trust value; W i3 The value is between W i2 and W i1 which means that the influence degree of change pattern Model3 is between Model2 and Model1, because it indicates that the electricity meter has been performing well; then, the overall trust value of each electricity meter is calculated using Equation (9):
[0114]
[0115] Specifically, multiply the weight Win i of each window by its exponential decay function value and accumulate them to obtain the overall trust value of this window; finally, take the average of the overall trust values of all windows, which is the final overall trust value; among them, the exponential decay formula is the penalty factor, l i is the number of abnormal data in each window, and different θ i are assigned to the three change patterns, denoted as θ i1 , θ i2 and θ i3 , satisfying:
[0116] θ i1 > θ i2 > θ i3 > 0; (10)
[0117] So that when there are continuous anomalies, the trust value can drop rapidly, when it is gradually recovering, the trust value will rise slowly, and when there are no anomalies, the trust value remains at a relatively high level, which can effectively reflect the trend of trust value changes under different change patterns;
[0118] (3 - 3) Comprehensive trust value calculation:
[0119] The calculation of the comprehensive trust value is the weighted sum of the direct trust value and the overall trust value to obtain the final trust value of each smart electricity meter; the comprehensive trust value not only considers the direct trust value but also combines the overall trust value. Among them, the direct trust value reflects the abnormal conditions of the smart electricity meter in different time periods; the overall trust value is based on the analysis of the change pattern of abnormal data in the sliding window, reflecting the comprehensive performance of the smart electricity meter in terms of continuous stability, recovery ability, etc., and can comprehensively evaluate the trust value of the smart electricity meter; the calculation of the comprehensive trust value is as follows:
[0120] TVi = λ × DT i + (1 - λ) × OT i , (11)
[0121] Where λ is a parameter between 0 and 1, used to balance the influence of direct trust value and overall trust value on the comprehensive trust value;
[0122] (4) Optimal secure path selection:
[0123] (4-1) Each smart meter SM i uses itself as the starting path node to construct the optimal secure path p i ; Let i = 1, and the starting path node is x1 = SM i , that is, the optimal secure path is p i = x1;
[0124] (4-2) Determine whether there is a data concentrator DC in the neighbor set N i of the current node; if it exists, use DC as the next-hop path node and add it to the optimal secure path p i , that is, p i = p i ∪ {DC}, and the path selection ends;
[0125] (4-3) If there is no data concentrator in the neighbor set, traverse the nodes in the set N i , where the comprehensive trust value of each neighbor node SM j is represented as TV j , select the node with the maximum comprehensive trust value TV j that is not in the current path p i as the next-hop path node x i+1 , and add it to the path p i , that is, p i = p i ∪ {x i+1}, x i+1 is represented as:
[0126]
[0127] Then, filter out the nodes that already exist in the path from the neighbor node set, that is, N i = N i / {x i+1}, update i = i + 1, and return to (4-2).
[0128] The present invention constructs a power network topology containing 10 smart meters (numbered SM0 to SM9 respectively) to simulate data transmission. The neighbor relationships between nodes are defined as follows (the direct neighbor nodes of the node are in parentheses): N0 = [SM1, SM2, SM3], N1 = [SM0, SM2, SM3, SM4, SM5], N2 = [SM0, SM1, SM3, SM4, SM5], N3 = [SM0, SM1, SM2, SM4, SM5], N4 = [SM1, SM2, SM3, SM6, SM7, SM8], N5 = [SM1, SM2, SM3, SM6, SM7, SM8], N6 = [SM4, SM5, SM7, SM8, SM9], N7 = [SM4, SM5, SM6, SM8, SM9], N8 = [SM4, SM5, SM6, SM7, SM9], N9 = [SM6, SM7, SM8]. Each meter energy dataset contains energy data every 30 minutes (i.e., cycle t = 30 min) for 100 days, including timestamp (Timestamp i_t ), hour (Hour i_t ), minute (Minute i_t ), temperature (Temperature i_t ), "whether someone is at home" (Occupier i_t ) and meter reading (Data i_t ). Then each group of experiments has 48010 pieces of data (because each smart meter will have one more piece of data at 0:00 on the 101st day, so each group has 10 more pieces of data). The data generation is divided into the following steps:
[0129] First, generate a timestamp sequence covering from December 1, 2023 to March 10, 2024, with an interval of 30 minutes, in the form of YYYY-MM-DD HH:MM, where HH and MM form the hour and minute respectively;
[0130] Then, according to the date represented by each timestamp YYYY-MM-DD, as well as the hour and minute, combined with the weather conditions and work patterns in the northern winter, generate the temperature (unit: °C, value less than 10) and the label of "whether someone is at home" (taking values of 0 or 1, 0 means no one is at home, 1 means someone is at home) in sequence;
[0131] Finally, based on the peak electricity consumption period (from 7 am to 9 am or from 6 pm to 10 pm), the normal electricity consumption period (from 9 am to 6 pm) and the low - valley electricity consumption period (other times), and comprehensively considering the influence of temperature and "whether someone is at home", generate the meter reading Data i_t . The specific generation rules are as follows:
[0132] During peak electricity consumption periods, if Occupier i_t = 1, then Data i_t ∈(3, 5); if Occupier i_t = 0, then Data i_t ∈(1, 2). Regardless of whether Occupier i_t = 1 or Occupier i_t = 0, the smaller the value of Temperature i_t , the larger the value of Data i_t .
[0133] During normal electricity consumption periods, if Occupier i_t = 1, then Data i_t ∈(2, 4); if Occupier i_t = 0, then Data i_t ∈(1, 2). Regardless of whether Occupier i_t = 1 or Occupier i_t = 0, the smaller the value of Temperature i_t , the larger the value of Data i_t .
[0134] During off-peak electricity consumption periods, if Occupier i_t = 1, then Data i_t ∈(1, 3); if Occupier i_t = 0, then Data i_t ∈(1, 2). Regardless of whether Occupier i_t = 1 or Occupier i_t = 0, the smaller the value of Temperature i_t , the larger the value of Data i_t .
[0135] The generated dataset finally includes the timestamp (Timestamp i_t ), hour (Hour i_t ), minute (Minute i_t ), temperature (Temperature i_t ), "whether someone is at home" (Occupier i_t ) and the electricity meter reading (Data i_t ).
[0136] Four experiments are set up in the present invention, which are as follows:
[0137] Experiment 1: It includes 500 groups of random experiments. The optimal safe paths and path safety values generated by the method of the present invention are used to verify that the node with the highest trust value can be selected as the forwarding node at each hop, so as to verify the effectiveness of the method of the present invention.
[0138] Experiment 2: It includes 500 groups of random experiments. The method of the present invention is compared with the random path selection method not based on trust by using the path safety value, that is, the sum of the comprehensive trust values of all nodes forming the path as the measurement standard, so as to verify the efficiency of the method of the present invention.
[0139] Experiment 3: 500 groups of random experiments are carried out. Taking the false negative rate, false positive rate and F1 score as the standards, the accuracy of the Isolation Forest algorithm iForest in the method of the present invention when dealing with abnormal data with different abnormal ratios (10%-40%) is evaluated.
[0140] Experiment 4: Taking the false negative rate, false positive rate and F1 score as the standards, the accuracy of the Isolation Forest algorithm iForest in the method of the present invention is compared with that of the traditional anomaly detection models Local Outlier Factor (LOF) and Support Vector Machine (SVM).
[0141] Figures 2 - 7 The optimal safe path diagrams of the 88th, 104th, 122nd, 245th, 367th and 421st groups of Experiment 1 are shown. It can be seen from the figure that the six path safety values generated by the method of the present invention are 1.923, 1.937, 1.945, 1.841, 1.859 and 1.914 respectively, and the path safety values of the optimal safe paths of this method are all the maximum values in the current network topology.
[0142] Figure 8 The comparison of the path safety values of the optimal safe paths selected by the method of the present invention and the random path selection method not based on trust in 500 groups of Experiment 2 is shown. As shown in the figure, the safety values of the optimal safe paths selected by the present invention are always greater than or equal to the path safety values of the random path selection method. For example, in the 88th group, the path safety value of the method of the present invention is 1.86914, while the path safety value of the random path selection method is 1.84956; in the 104th group, the path safety value of the method of the present invention is 1.86893, while the path safety value of the random path selection method is 1.83897; in the 167th group, the path safety value of the method of the present invention is 1.87516, while the path safety value of the random path selection method is 1.81975.
[0143] Figure 9Shows the average false negative rate, false positive rate, and F1 score of the Isolation Forest algorithm iForest used in the present invention in the detection of electricity meter energy data with different abnormal ratios in Experiment 3. Among them, the abnormal ratio of each electricity meter energy data is [0.1, 0.15, 0.2, 0.25, 0.3, 0.35, 0.4, 0.3, 0.2, 0.1] in turn. It can be seen from the figure that the Isolation Forest algorithm iForest has a high F1 score, a low false positive rate, and a low false negative rate under different abnormal ratios of electricity meter energy data. The average value of the F1 score of iForest is [97.35%, 96.49%, 96.53%, 96.69%, 96.80%, 96.75%, 96.82%, 96.74%, 96.69%, 96.73%], the average value of FPR is [11.38%, 11.93%, 11.63%, 11.47%, 11.32%, 11.68%, 11.67%, 11.79%, 11.99%, 11.95%], and the average value of FNR is [4.19%, 2.60%, 4.17%, 3.88%, 3.67%, 3.70%, 3.57%, 3.69%, 3.73%, 3.66%].
[0144] Figure 10 Shows the false negative rate of the Isolation Forest algorithm iForest, LOF, and SVM in Experiment 4. It can be seen from the figure that the false negative rate of the Isolation Forest algorithm iForest is the lowest, with an average false negative rate of about 4.53% and the lowest false negative rate of 2.79%, showing good anomaly detection ability, while the false negative rates of the other two algorithms are significantly higher. Among them, the average false negative rate of LOF is 20.49%, and the highest false negative rate is as high as 21.54%. SVM is between the Isolation Forest algorithm iForest and LOF, but the average false negative rate is also relatively high at 15.34%.
[0145] Figure 11 Shows the false positive rate of the Isolation Forest algorithm iForest, LOF, and SVM in Experiment 4. It can be seen from the figure that there is an obvious gap among the three algorithms. Among them, the false positive rate of LOF is much higher than that of the other two algorithms, with an average false positive rate of 82.47%, indicating that LOF has the worst ability to identify normal data. Similarly, SVM also has a high false positive rate, with an average false positive rate of 60.37%. While the average false positive rate of the Isolation Forest algorithm iForest is only 18.80%, and its performance in distinguishing normal data and abnormal data is the best.
[0146] Figure 12Shows the F1 scores of the Isolation Forest algorithm iForest, LOF, and SVM in Experiment 4. As can be seen from the figure, the Isolation Forest algorithm iForest has the highest average F1 score, averaging 95.55%, showing consistent and excellent anomaly detection performance. The F1 score of LOF is the lowest, averaging 79.44%, with the worst detection performance. SVM is between the other two, with an average F1 score of 84.91%.
Claims
1. A trust-based data integrity attack defense method for forwarding nodes in an advanced measurement system. Its characteristics include the following steps: (1) Data collection: Each user i, i=1,2,...,m has a smart meter SM i ,i=1,2,...,m, where m is the total number of users; each smart meter SM i Collect user data Message regularly in period t i_t , including communication data Traffic i_t , Forward data i_t and Energy Data i_t ,Energy i_t Including hours i_t 、Minute i_t Temperature i_t , "Is anyone home?" i_t and meter readings Data i_t ; In addition, each smart meter SM i You can also obtain each neighborhood smart meter SM from the power company. j ,j=1,2,...,n i Meter energy data, where n i Smart Meter SM i Number of smart meters in the neighborhood; SM i The neighbor smart meter refers to the smart meter SM in the same network topology. i Other smart meters reachable in one hop; each smart meter SM i n i Neighborhood Smart MetersSM j ,j=1,2,...,n i Constitute its neighbor node set N i ; (2) Anomaly Detection: (2-1) Forming a data set: Each smart meter SM i Get the SM of each neighbor node from the power company j ,j=1,2,...,n i Energy data for the previous d cycles, denoted as Energy j_c ,c=td,t-d+1,...,t-1, forming data set D j ,j=1,2,...,n i ; (2-2) Generate an isolated tree: Each smart meter SM i For each neighbor node SM j ,j=1,2,...,n i The electric meter energy dataset D j Perform sub-collection, SM i Randomly select ψ j ,ψ j ∈D j Get the sub-dataset D j '; Then, in the subdataset D j 'Randomly select feature set Q j ={Hour j_t ,Minute j_t ,Temperature j_t ,Occupier j_t ,Data j_t A feature q in j ; Then randomly select a value s within the range of the feature j As the split point, the data with eigenvalues less than this value is taken as the left child, otherwise it is the right child; repeat this process until each isolated tree generated reaches the specified depth h j_max Or the number of samples contained in each leaf node does not exceed the threshold n j_min ; (2-3) Constructing an isolation forest: Repeat the process of splitting the subtree described in (2-2) to obtain the energy data set D of the electric meter based on the neighboring nodes. j Step by step, build iTree j_1 ~iTree j_jnum Isolation tree, where jnum represents the number of iTree; finally, all constructed isolated trees are merged to form a complete isolation forest iForest j ; (2-4) Calculate the anomaly score: Calculate the Energy of each data in the isolation forest j_c The average path length is as shown in formula (1): Among them, H j (Energy j_c ) is dataEnergy j_t The average path length of the jth tree, E(H(Energy j_c ) is dataEnergy j_c The average path length is then normalized to obtain the normalized average path length C(d), as shown in formula (2): Where d is the meter energy dataset D j The number of data in H d is the harmonic number that provides the standardized benchmark; finally, substitute it into formula (3) to obtain the data Energy j_c Anomaly score: Among them, S (Energy j_c ,d) is the energy data of the electricity meter j_c ,Energy j_c ∈D j The abnormality score of is in the range of [0-1]; (2-5) Determine abnormal data: According to the abnormal score S(Energy j_c ,d), by setting the threshold θ j ,j=1,2,...n i Judgment DataEnergy j_c Is it abnormal data? j_c ,d)≥θ j When, dataEnergy j_c It is judged as abnormal data. When S(Energy j_c ,d)<θ j When, dataEnergy j_c It is judged as normal data; (3) Trust Assessment: (3-1) Direct trust value calculation: First, considering the difference in electricity consumption in different time periods of the day, a day is divided into three time periods: peak, off_peak, and low; secondly, D j The number of abnormal data in each time period and the total number of abnormal data in the data are recorded as N peak 、N off_peak 、N low 、N total , and then calculate the probability P1, P2, P3 of abnormal data occurring during peak power consumption period, flat power consumption period and valley power consumption period according to formula (4): Next, based on P1, P2, and P3 obtained from formula (4), a weighted sum is performed to obtain the direct trust value DT i , as shown in formula (5): DT i =P1×α+P2×β+P3×γ, (5) Among them, α, β, and γ are the weights of the three time periods; Specifically, α, β, and γ are calculated by initializing weights α1, β2, and γ3 through formula (6): Among them, weight_func is the weight adjustment function, as shown in formula (7): weight_func(p)=1-P; (7) (3-2) Calculation of overall trust value: First, obtain D j All d data in the data set; then, set a sliding window Window, the size of the sliding window Window length Set to w, step size Window λ is also set to w, then there are windows, and count the number of abnormal data in each window l; Window length =Window λ The window size and step length are matched with the temporal characteristics of the data, which can strike a balance between computational complexity and anomaly detection accuracy. Secondly, the threshold μ is set for each window. j ,j=1,2,3...,o, when the number of abnormal data l in the window is greater than the threshold μ, it is marked as change mode Model1: "smart meter with continuous abnormality", indicating that the meter is in an untrustworthy state for a long time; if the number of abnormal data l is greater than 0 but not greater than μ, it is marked as change mode Model2: "smart meter with gradual recovery", indicating that the meter is recovering to normal; if the number of abnormal data l is 0, it is marked as change mode Model3: "smart meter with high stability", indicating that the meter has been performing well; again, different weights W are assigned to the three change modes i , respectively W i1 , W i2 and W i3 ,satisfy: Then, the overall trust value of each meter is calculated using formula (9): Specifically, the weight Win of each window i With its exponential decay function value Multiply and accumulate to get the overall trust value of the window; finally, take the average of the overall trust values of all windows, which is the final overall trust value; where the exponential decay formula is the penalty factor, l i The number of abnormal data in each window is given different θ for the three change modes i , denoted as θ i1 ,θ i2 and θ i3 ,satisfy: i i1 >θ i2 >θ i3 >0; (10) (3-3) Comprehensive trust value calculation: The calculation of the comprehensive trust value is the weighted sum of the direct trust value and the overall trust value, so as to obtain the final trust value of each smart meter; the calculation of the comprehensive trust value is as follows: TV i =λ×DT i +(1-λ)×OT i , (11) Among them, λ is a parameter between 0 and 1, which is used to balance the impact of direct trust value and overall trust value on the comprehensive trust value; (4) Optimal safety path selection: (4-1) Each smart meter SM i Use yourself as the starting path node to build the optimal safe path p i ; Let i = 1, the starting path node is x1 = SM i , that is, the optimal safe path is p i =x1; (4-2) Determine the neighbor set N of the current node i Is there a data concentrator DC in the network? If so, DC is used as the next hop path node and added to the optimal security path p i In, that is, p i =p i ∪{DC}, path selection ends; (4-3) If there is no data concentrator in the neighbor set, traverse the set N i , where each neighbor node SM j The comprehensive trust value is expressed as TV j , select the comprehensive trust value TV j The largest one that is not on the current path p i The node in the path is the next hop node x i+1 , and add to path p i In, that is, p i =p i ∪{x i+1 }, x i+1 It is expressed as: Then, the nodes that already exist in the path are filtered out from the neighbor node set, that is, N i =N i / {x i+1 }, update i=i+1, and return to (4-2).