An industrial internet of things abnormal flow online identification method
Patent Information
- Application Number
- CN202511753422.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2045-11-26
AI Technical Summary
[0004]有益效果:本发明首先获取目标叶节点,然后根据目标叶节点的最近邻向量序列中的各向量的最优属性以及当前时刻与当前时刻的前一时刻下目标叶节点的最优属性增益差和次优属性增益差,得到当前时刻下目标叶节点的置信度表征值,之后判断置信度表征值是否大于预设置信阈值,若是,则根据Hoeffding不等式对当前时刻下的目标叶节点进行分裂决策,否则,则根据到达目标叶节点的所有向量以及当前网络性能指标值向量和最近邻向量序列所形成的集合中不同类别向量的数量差,得到当前时刻下目标叶节点的分裂必要性指标值,根据分裂必要性指标值对当前时刻下的目标叶节点进行分裂决策,并输出当前时刻工业物联网异常流量的识别结果。且本发明在进行分裂决策时加入分裂必要性指标值,可提高分裂决策的正确性、可靠性、及时性,进而能够提高对工业物联网异常流量识别的准确性和可靠性。
Smart Images

Figure CN121509005B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of abnormal traffic identification technology, and more specifically to an online method for identifying abnormal traffic in the industrial Internet of Things. Background Technology
[0002] Currently, to ensure production safety, prevent equipment failures, and guard against cyberattacks, abnormal traffic identification is commonly performed on the Industrial Internet of Things (IIoT). Existing technologies typically use the Hoeffding tree algorithm for this purpose. However, current methods for identifying abnormal traffic in the IIoT often rely on the Hoeffding inequality to determine when to split the leaf nodes of the decision tree. In other words, the existing leaf node splitting strategy is usually dominated by the Hoeffding inequality. That is, after a new sample is input into the Hoeffding tree and reaches a certain leaf node, the splitting is then performed by calculating the current leaf node's statistics and the Hoeffding inequality. Hoeffding's inequality is used to output splitting decisions. However, traffic data in the Industrial Internet of Things (IIoT) is often non-stationary, bursty, and heterogeneous. Hoeffding's inequality relies on assumptions such as independent and identically distributed data and stable distribution. Therefore, when using Hoeffding's inequality to determine when to split leaf nodes, incorrect splitting decisions may occur, leading to splitting delays or incorrect splitting. This results in misidentification or missed identification when identifying abnormal traffic in the IIoT. Therefore, improving the correctness, reliability, and timeliness of splitting decisions to enhance the accuracy and reliability of abnormal traffic identification in the IIoT has become an urgent problem to be solved. Summary of the Invention
[0003] To address the aforementioned problems, this invention provides an online method for identifying abnormal traffic in the Industrial Internet of Things (IIoT), the specific technical solution of which is as follows: One embodiment of the present invention provides a method for online identification of abnormal traffic in the industrial Internet of Things, comprising the following steps: Obtain the target leaf node, which is the leaf node reached by the current network performance index value vector after it is input into the Hofding tree, and the current network performance index value vector is the vector of the Industrial Internet of Things at the current moment; Based on the optimal attributes of each vector in the nearest neighbor vector sequence of the target leaf node, as well as the difference between the optimal attribute gain and the difference between the second-best attribute gain of the target leaf node at the current time and the time before the current time, the confidence characterization value of the target leaf node at the current time is obtained. If the confidence level is greater than a preset confidence threshold, a splitting decision is made for the target leaf node at the current moment according to the Hoeffding inequality. Otherwise, the splitting necessity index value of the target leaf node at the current moment is obtained based on the difference in the number of different categories of vectors in the set formed by all vectors reaching the target leaf node, the current network performance index value vector, and the nearest neighbor vector sequence. A splitting decision is made for the target leaf node at the current moment based on the splitting necessity index value, and the identification result of abnormal traffic in the industrial IoT at the current moment is output.
[0004] Beneficial Effects: This invention first obtains the target leaf node. Then, based on the optimal attributes of each vector in the nearest neighbor vector sequence of the target leaf node, and the difference between the optimal and second-best attribute gain of the target leaf node at the current time and the time before the current time, it obtains the confidence characterization value of the target leaf node at the current time. Next, it determines whether the confidence characterization value is greater than a preset confidence threshold. If so, it makes a splitting decision for the target leaf node at the current time according to the Hoeffding inequality. Otherwise, it obtains the splitting necessity index value of the target leaf node at the current time based on the difference in the number of different categories of vectors in the set formed by all vectors reaching the target leaf node, the current network performance index value vector, and the nearest neighbor vector sequence. It then makes a splitting decision for the target leaf node at the current time based on the splitting necessity index value and outputs the identification result of abnormal traffic in the Industrial Internet of Things (IIoT) at the current time. Furthermore, by incorporating the splitting necessity index value into the splitting decision process, this invention improves the correctness, reliability, and timeliness of the splitting decision, thereby enhancing the accuracy and reliability of identifying abnormal traffic in the IIoT. Attached Figure Description
[0005] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0006] Figure 1 This is a flowchart of an online method for identifying abnormal traffic in the industrial Internet of Things according to the present invention. Detailed Implementation
[0007] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention are within the protection scope of the embodiments of the present invention.
[0008] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.
[0009] This embodiment provides a method for online identification of abnormal traffic in the industrial Internet of Things (IoT), which is described in detail below: like Figure 1 As shown, this online method for identifying abnormal traffic in the Industrial Internet of Things includes the following steps: Step S001: Obtain the target leaf node. The target leaf node is the leaf node reached by the current network performance index value vector after it is input into the Hofding tree. The current network performance index value vector is the vector of the Industrial Internet of Things at the current moment.
[0010] This embodiment first collects network performance index values of different attribute types corresponding to the Industrial Internet of Things (IIoT) at any given time. The set of network performance index values of all attribute types corresponding to the IIoT at any given time is denoted as the network performance index value vector of the IIoT at that time. The network performance index value vector of the IIoT at the current time is denoted as the current network performance index value vector of the IIoT at the current time. The attribute types of the network performance index values obtained at each time in this embodiment include, but are not limited to, average packet length, packet rate per second, and protocol distribution ratio. Average packet length, packet rate per second, and protocol distribution ratio are all traffic data. Index values with the same attribute have the same position in different vectors. The average packet length in different vectors is the same attribute, the packet rate per second in different vectors is the same attribute, and the protocol distribution ratio in different vectors is the same attribute. The process of obtaining the average packet length, packet rate per second, and protocol distribution ratio of the IIoT at any given time is a known technology. Furthermore, the network performance index values such as average packet length, packet rate per second, and protocol distribution ratio in this embodiment are index values after removing the dimensions. For example, if the network performance metric vector at a certain moment in this embodiment is composed of average packet length, packet rate per second, and protocol distribution ratio, and if the average packet length at that moment is 123.33 bytes, the packet rate per second is 5 packets per second, the TCP protocol ratio is 0.8, the UDP protocol ratio is 0.2, and the ICMP protocol ratio is 0.0, then the network performance metric vector at that moment would be {123.33,5,0.8,0.2,0}. The attribute types of the network performance metric vector at that moment include average packet length, packet rate per second, and protocol distribution ratio, with the TCP protocol ratio, UDP protocol ratio, and ICMP protocol ratio belonging to the protocol distribution ratio category.
[0011] After obtaining the current network performance index value vector of the Industrial Internet of Things (IIoT) at the current moment and the network performance index value vector of the IIoT at each moment, this embodiment obtains the target leaf node. The subsequent part of this embodiment is to analyze and obtain the splitting strategy of the target leaf node. The process of obtaining the target leaf node is as follows: after the current network performance index value vector is input into the Hofding tree, the leaf node reached by the current network performance index value vector is recorded as the target leaf node.
[0012] Therefore, the target leaf node was obtained through the above process in this embodiment.
[0013] Step S002: Based on the optimal attributes of each vector in the nearest neighbor vector sequence of the target leaf node, and the optimal attribute gain difference and the second-best attribute gain difference of the target leaf node at the current time and the time before the current time, obtain the confidence characterization value of the target leaf node at the current time.
[0014] Currently, the traditional Hoeffding tree algorithm is commonly used for abnormal traffic identification in the Industrial Internet of Things (IIoT). This algorithm is based on the Hoeffding inequality, which states there exists a boundary, or Hoeffding Bound. The Hoeffding Bound theoretically relies on assumptions such as independent and identically distributed (i.i.d.) data and stable distribution. When the gain difference between the optimal and second-best attributes exceeds this boundary value, a split condition is triggered. Both the optimal and second-best attributes can be calculated using the Hoeffding tree algorithm, which is a standard operation within the algorithm. However, in the field of abnormal traffic identification in the IIoT, these assumptions do not always hold. This is because traffic data in the IIoT typically exhibits characteristics such as non-stationarity, burstiness, and heterogeneity. These characteristics can undermine the assumptions of independent and identically distributed data and stable distribution. For example, network topology, equipment load, and communication protocol stack states change over time. Furthermore, differences in communication modes at different process stages and variations in the data rates of PLCs and sensors lead to changes in the proportion of abnormal samples, resulting in a lack of constant stationarity in the traffic data and potentially disrupting the Hoeffding inequality. The constraints of the bounds can be limited or even ineffective. Therefore, if the decision of when to split a leaf node is made solely based on the Hoeffding inequality, it may lead to incorrect splitting decisions, resulting in problems such as splitting delays or incorrect splitting. This can cause misidentification or missed identification when identifying abnormal traffic in the Industrial Internet of Things (IIoT). To reduce or avoid incorrect splitting decisions, this embodiment will subsequently perform adaptive splitting decisions based on features such as the obviousness of the splitting signal (or the confidence level of splitting according to the Hoeffding inequality), the distribution complexity of leaf nodes, and the drift degree of leaf node category distribution. This improves the correctness, reliability, and timeliness of splitting decisions, thereby improving the accuracy and reliability of identifying abnormal traffic in the IIoT. In other words, after inputting the current network performance index value vector into the Hoeffding tree, to improve the correctness, reliability, and timeliness of splitting decisions and ensure the accuracy and reliability of identifying abnormal traffic in the IIoT, the decision should not be made directly based on the Hoeffding inequality. The split determination based on the Hoeffding Bound, or rather, should not be based directly on the Hoeffding inequality. Instead, it should first determine whether the leaf node to which the input current network performance index value vector belongs satisfies the Hoeffding Bound theory, or whether the Hoeffding inequality results in constraint failure or insufficient constraint on the leaf node to which the current network performance index value vector belongs. In other words, this embodiment will first determine the optimal attribute of each vector in the nearest neighbor vector sequence of the target leaf node, as well as the optimal attribute gain difference and the second-best attribute gain difference of the target leaf node at the current time and the time before the current time.The confidence level of the target leaf node at the current time is obtained. This confidence level reflects whether the leaf node to which the current network performance index value vector belongs satisfies the Hoeffding Bound theory, or whether the Hoeffding inequality has a constraint failure or insufficient constraint on the target leaf node. The Hoeffding Bound is the core mathematical expression of the Hoeffding inequality, used to quantify the upper probability limit of the difference between the estimated and true values of statistics. Therefore, the specific process for obtaining the confidence level of the target leaf node at the current time in this embodiment is as follows: First, we obtain the nearest neighbor vector sequence of the target leaf node at the current time, the optimal attribute of each network performance index value vector in the nearest neighbor vector sequence, and the information gain of the optimal attribute and the information gain of the second-best attribute of the target leaf node at different times. The process of obtaining the information gain of the optimal attribute and the information gain of the second-best attribute of the target leaf node at different times is a well-known technique.
[0015] The process of obtaining the nearest neighbor vector sequence of the target leaf node at the current moment is as follows: The set of all network performance index value vectors received by the target leaf node up to the current moment is denoted as the comprehensive vector set corresponding to the target leaf node at the current moment. The comprehensive vector set includes the current network performance index value vector. The set of vectors in the comprehensive vector set other than the current network performance index value vector is denoted as the remaining set. The vectors received by the target leaf node are the vectors that have arrived at the target leaf node. All vectors in the remaining set are sorted according to the order of vector input or the order in which the vectors arrive at the target leaf node. The sorting result is denoted as the temporal vector sequence of the target leaf node at the current moment. The sequence of the last preset number of network performance index value vectors in the temporal vector sequence is denoted as the nearest neighbor vector sequence of the target leaf node at the current moment. The vectors in the nearest neighbor vector sequence are arranged according to the order of vector input or the order in which the vectors arrive at the target leaf node. In specific applications, the implementer needs to set the preset number value according to the actual situation. For example, in this embodiment, the preset number can be set to 5.
[0016] The optimal attribute of any network performance metric vector is the feature or attribute that, when input into a Hofding tree, most effectively distinguishes between normal and abnormal traffic states at the current leaf node. Here, the current leaf node is the actual leaf node reached by the network performance metric vector. For example, assuming the average packet length in the network performance vector is 256 bytes, the packet rate is 1800 packets per second, TCP accounts for 45%, UDP accounts for 40%, and other protocols account for 15%, calculating the information gain of each attribute yields an information gain of 0.24 for the average packet length and 0.19 for the packet rate per second. The information gain of the protocol distribution ratio is 0.43. Since the information gain of the protocol distribution ratio is the largest, it is the optimal attribute of the network performance vector. This also means that the protocol distribution ratio is the most critical indicator for identifying abnormal traffic. The optimal attribute is selected by comparing the information gain or Gini index of each attribute. The calculation process of information gain is well known. In the Hoeffding tree, each input vector starts from the root node and flows down layer by layer according to the attribute value, eventually stopping at a certain leaf node. This leaf node is the current leaf node of the vector and is also the node for split judgment and classification decision.
[0017] In a Hoeffding tree, the optimal attribute of a leaf node at a given time refers to the feature attribute that maximizes classification accuracy when the input vector received by that leaf node arrives at that time. However, the optimal attribute of a leaf node at a given time is different from the optimal attribute of the vector arriving at that leaf node at a given time. Specifically: the optimal attribute of a node refers to the globally optimal splitting feature calculated based on all historical accumulated data. For example, if a leaf node has processed 1000 samples or vectors, its optimal attribute might be the protocol distribution ratio, which is a statistical result of all data. In other words, the optimal attribute of a node is global and depends on long-term data statistics. The optimal attribute of a vector, on the other hand, is the locally optimal classification attribute of that vector at the leaf node it arrives at. For example, when a vector arrives at leaf node A, the packet rate per second might be chosen as the optimal attribute based on local data. However, this is only for that vector; the optimal attribute of a vector is instantaneous and only reflects the local features of the current sample. Obtaining the optimal attributes of leaf nodes and vectors is a known technique.
[0018] After obtaining the nearest neighbor vector sequence, the optimal attributes of each network performance index value vector in the nearest neighbor vector sequence, and the information gain of the optimal and suboptimal attributes of the target leaf node at different times, this embodiment then obtains the flip sequence of the target leaf node at the current time based on the optimal attributes of each network performance index value vector in the nearest neighbor vector sequence of the target leaf node at the current time. The time before the current time is denoted as time t-1, and the current time is time t. Then, the information gain of the optimal attribute of the target leaf node at the current time is calculated minus the information gain of the optimal attribute of the target leaf node at time t-1, and the result is denoted as the optimal attribute gain difference. The information gain of the suboptimal attribute of the target leaf node at the current time is calculated minus the information gain of the suboptimal attribute of the target leaf node at time t-1, and the result is denoted as the suboptimal attribute gain difference. Finally, based on the flip sequence of the target leaf node at the current time, the optimal attribute gain difference, and the suboptimal attribute gain difference, the confidence characterization value of the target leaf node at the current time is obtained.
[0019] Based on the optimal attributes of each network performance index value vector in the nearest neighbor vector sequence of the target leaf node at the current time, the specific process of obtaining the flip sequence of the target leaf node at the current time is as follows: Based on the differences or similarities in the optimal attributes of adjacent network performance index value vectors in the nearest neighbor vector sequence, a flip sequence for the target leaf node at the current time is constructed. If the optimal attribute of the a-th network performance index value vector in the nearest neighbor vector sequence is the same as the optimal attribute of the (a+1)-th network performance index value vector in the nearest neighbor vector sequence, then 0 is assigned to the a-th parameter in the flip sequence, meaning the value of the a-th parameter in the flip sequence is 0. If the optimal attribute of the a-th network performance index value vector in the nearest neighbor vector sequence is different from the optimal attribute of the (a+1)-th network performance index value vector, then 1 is assigned to the a-th parameter in the flip sequence, meaning the value of the a-th parameter in the flip sequence is 1. For example, if the total number of vectors in the nearest neighbor vector sequence is 5, and the optimal attribute of the first vector in the nearest neighbor vector sequence is... The optimal attributes of the first and second vectors are: average packet length, average packet length, packet rate per second, protocol distribution ratio, and protocol distribution ratio. The optimal attributes of the first and second vectors are the same (average packet length). Therefore, the value of the first parameter in the flip sequence is 0. The optimal attributes of the second and third vectors are different, so the value of the second parameter in the flip sequence is 1. The optimal attributes of the third and fourth vectors are also different, so the value of the third parameter in the flip sequence is 1. The optimal attributes of the fourth and fifth vectors are the same (protocol distribution ratio), so the value of the fourth parameter in the flip sequence is 0. Therefore, the flip sequence is {0,1,1,0}.
[0020] The specific process for obtaining the confidence representation value of the target leaf node at the current time based on the flip sequence of the target leaf node, the optimal attribute gain difference, and the second-best attribute gain difference is as follows: Calculate the mean of all parameter values in the flip sequence and denote it as the average flip characteristic value. Perform a negative correlation mapping on the average flip characteristic value and denote the mapping result as the first characterization value. Here, a negative exponential function with a constant e as the base is used for mapping. The smaller the average flip characteristic value, the more consistent the optimal attribute of the vector received by the target leaf node is over a long period, and the more stable the splitting signal is. At this time, the target leaf node is more in line with or satisfies the Hoeffding Bound theory, and the more reliable and effective the Hoeffding Bound constraint is. At this time, the splitting determination of the target leaf node at the current moment based on Hoeffding Bound or Hoeffding inequality is more reliable and trustworthy. Obtain the optimal attribute gain change rate and the second-best attribute gain change rate. The optimal attribute gain change rate is... The rate of change of suboptimal attribute gain is , The optimal attribute gain difference. The difference in gain between suboptimal attributes is represented by `max()`, which is the function to find the maximum value. Let be the information gain of the optimal attribute of the target leaf node at time t-1. Let be the information gain of the suboptimal attribute of the target leaf node at time t-1. The preset constant is used to prevent the denominator from being 0, thus ensuring the calculation is valid. In specific applications, the implementer needs to set the value of the preset constant according to the actual situation, but the preset constant must be greater than 0. For example, in this embodiment, the preset constant value can be 1. Calculate the absolute value of the difference between the optimal attribute gain rate of change and the second-best attribute gain rate of change, and record it as the rate of change difference. Normalize the rate of change difference, and record the result of the normalization as the second characterization value. Here, the hyperbolic tangent function is used to normalize the rate of change difference. The rate of change difference represents the consistency between the optimal attribute rate of change and the second-best attribute rate of change. The more obvious the difference, the worse the consistency of change, indicating that a certain attribute has a significant advantage or the fluctuation trend is significantly different, and it is better to distinguish the split direction. At this time, it is more reliable and trustworthy to use the Hoeffding Bound or Hoeffding inequality to determine the split of the target leaf node at the current moment. Calculate the weighted sum of the first characterization value and the second characterization value, and use it as the confidence characterization value of the target leaf node at the current moment. The specific calculation expression of the confidence characterization value of the target leaf node at the current moment is: in, This represents the confidence level of the target leaf node at the current moment. As the first weight value, Here, exp() is the second weight value, and it is an exponential function with base e. Here, is the value of the nth parameter in the flipped sequence, N is the total number of parameters in the flipped sequence, and tanh() is the hyperbolic tangent function. In specific applications, the implementer needs to set the values of the second weight value and the first weight value according to the actual situation. However, Hoeffding Bound is essentially a statistical inference, so the second representation value is more important to the confidence representation value. Therefore, this embodiment requires the second weight value to be greater than the first weight value. (The last part, "can be set," appears to be an error and can be left untranslated.) , And because smaller and The larger the value, the more reliable and trustworthy the splitting determination of the target leaf node at the current time is based on the Hoeffding Bound or Hoeffding inequality. smaller and When it is larger, The larger, therefore when A larger value indicates that the optimal attribute of the target leaf node switches infrequently, the splitting signal of the target leaf node is more stable, the difference between the rate of change of the optimal attribute and the rate of change of the second-best attribute is more obvious, and the Hoeffding Bound constraint is more reliable. This means that using the Hoeffding Bound or Hoeffding inequality to determine the splitting of the target leaf node at the current moment is more reliable and trustworthy. Conversely, a smaller value indicates a less frequent or infrequent splitting of the target leaf node. The smaller the value, the more frequently the optimal attribute of the target leaf node switches, the unstable splitting signal of the target leaf node, the less distinct the difference between the rate of change of the optimal attribute and the rate of change of the second-best attribute, and the worse the effect of the Hoeffding Bound constraint. This indicates that it is less reliable and untrustworthy to make splitting decisions for the target leaf node at the current moment based on Hoeffding Bound or Hoeffding inequality. Higher reliability and trustworthiness indicate that subsequent splitting decisions for the target leaf node at the current moment based on Hoeffding Bound or Hoeffding inequality are less likely to result in incorrect splitting decisions or problems such as splitting delay or incorrect splitting. Lower reliability and trustworthiness indicate that subsequent splitting decisions for the target leaf node at the current moment based on Hoeffding Bound or Hoeffding inequality are more likely to result in incorrect splitting decisions or problems such as splitting delay or incorrect splitting.
[0021] Therefore, this embodiment can obtain the confidence level characterization value of the target leaf node at the current time through the above process.
[0022] Step S003: Determine whether the confidence level is greater than a preset confidence threshold. If so, make a splitting decision for the target leaf node at the current moment according to the Hoeffding inequality. Otherwise, obtain the splitting necessity index value of the target leaf node at the current moment based on the difference in the number of different categories of vectors in the set formed by all vectors reaching the target leaf node, the current network performance index value vector, and the nearest neighbor vector sequence. Make a splitting decision for the target leaf node at the current moment based on the splitting necessity index value and output the identification result of abnormal traffic of the industrial IoT at the current moment.
[0023] Since the confidence level of the target leaf node at the current moment reflects the reliability and trustworthiness of splitting the target leaf node based on Hoeffding Bound or Hoeffding inequality, this embodiment will next determine whether to use Hoeffding Bound or Hoeffding inequality for splitting based on the obtained confidence level. That is, it will determine whether the confidence level of the target leaf node at the current moment is greater than a preset confidence threshold. If so, it indicates that splitting the target leaf node at the current moment based on Hoeffding Bound or Hoeffding inequality is more trustworthy and reliable, and is less likely to cause problems such as incorrect splitting decisions, splitting delays, or incorrect splitting. Therefore, when the confidence level is greater than the preset confidence threshold, the target leaf node at the current moment can be split based on Hoeffding Bound or Hoeffding inequality. In specific applications, implementers need to set the preset confidence threshold according to the range of confidence level values, experimental statistics, and other actual conditions. For example, the median of the confidence level range, 0.5, can be selected as the preset confidence threshold.
[0024] When the confidence value of the target leaf node at the current moment is not greater than the preset confidence threshold, it indicates that making a splitting decision for the target leaf node at the current moment based on Hoeffding Bound or Hoeffding inequality is unreliable and prone to problems such as incorrect splitting decisions, splitting delays, or incorrect splitting. In order to avoid problems such as incorrect splitting decisions, splitting delays, or incorrect splitting, this embodiment obtains the splitting necessity index value of the target leaf node at the current moment based on the difference in the number of different categories of vectors in the new set formed by all vectors reaching the target leaf node, the current network performance index value vector, and the nearest neighbor vector sequence, and makes a splitting decision for the target leaf node at the current moment based on the splitting necessity index value of the target leaf node at the current moment.
[0025] The specific calculation process for the splitting necessity index value of the target leaf node at the current moment is as follows, based on the difference in the number of vectors of different categories in the set formed by all vectors reaching the target leaf node, the current network performance index value vector, and the nearest neighbor vector sequence: Since the fundamental purpose of leaf node splitting in a Hofding tree is to improve the algorithm's future prediction discrimination and generalization ability, and to maintain the homogeneity of vectors or samples within a leaf node as much as possible in the feature space, a leaf node should trigger a split when the sample set it receives shows significant internal structural differences or when samples within a leaf node no longer conform to a single pattern. This is to improve or ensure the algorithm's future prediction discrimination and generalization ability. In other words, when the leaf node distribution is complex and the distribution of samples within a corresponding leaf node shifts significantly over time with the arrival of new samples, it indicates that the discriminative ability of the leaf node has decreased, and the samples within the leaf node are becoming increasingly homogeneous. The homogeneity of some samples is also decreasing. At this time, it is necessary to re-divide the sample space by splitting to ensure or improve the future prediction discrimination and generalization ability of the algorithm. That is, when the distribution of leaf nodes is relatively mixed, it means that the discrimination ability of the nodes is decreasing. It is necessary to improve the purity of the class within the node by splitting, thereby enhancing the discrimination of future prediction. When the distribution of samples within the node changes significantly over time with the arrival of new samples and a new clustering structure appears, it means that the homogeneity of samples within the node is decreasing. At this time, it is necessary to re-divide the sample space by splitting to improve the generalization ability. In this embodiment, the sample refers to a vector.
[0026] Based on the above analysis, this embodiment first divides vectors of the same type in the comprehensive vector set into the same set, and denots them as subsets. The comprehensive vector set includes the current network performance index value vector and all network performance index value vectors received by the target leaf node at the current time. The vectors received by the leaf node refer to all vectors that reach the corresponding leaf node. All vectors in the same subset belong to the same category. Since the algorithm determines whether the corresponding vector belongs to the normal category or the abnormal category when a network performance index vector is input into the Hofding tree algorithm, it can be known that the number of subsets is at most two: one is the normal category subset and the other is the abnormal category subset. In other words, the vectors in the comprehensive vector set may belong to the normal category or the abnormal category.
[0027] Then, based on the differences in the same attribute parameters between different subsets and the number of vectors in different subsets, the distribution complexity of the target leaf node at the current time is obtained. The specific process for obtaining the distribution complexity of the target leaf node at the current time is as follows: First, a feature difference set is obtained based on the differences in the same attribute parameters between different subsets. The calculation process for the feature difference set is as follows: The subset where all vectors belong to the normal category is denoted as the normal category subset, and the subset where all vectors belong to the abnormal category is denoted as the abnormal category subset. The mean of all vectors in the normal category subset is denoted as the normal mean vector, and the mean of all vectors in the abnormal category subset is denoted as the abnormal mean vector. The parameter value with the attribute of average packet length in the normal mean vector refers to the mean of the parameter values with the attribute of average packet length in all vectors in the normal category subset, and the same applies to the parameter values in the abnormal mean vector. The absolute value of the difference in the same attribute parameters between the normal mean vector and the abnormal mean vector is calculated, and the set of the absolute values of the differences in the same attribute parameters between the normal mean vector and the abnormal mean vector is denoted as the feature difference set. If the b-th parameter in the normal mean vector has the same attribute as the b-th parameter in the abnormal mean vector, then the b-th feature difference in the feature difference set is equal to the value of the b-th parameter in the normal mean vector. The absolute value of the difference between the value of the b-th parameter and the value of the b-th parameter in the outlier mean vector is calculated. Then, the subset with the largest number of vectors is obtained, and the number of vectors in this subset is recorded as the first quantity value. The subset with the smallest number of vectors is obtained, and the number of vectors in this subset is recorded as the second quantity value. The first quantity value is subtracted from the second quantity value, and this result is recorded as the quantity difference. The result of normalizing the quantity difference and then performing a negative correlation mapping is recorded as the first eigenvalue. Here, normalization is achieved using the total number of vectors in the comprehensive vector set, and negative correlation mapping is achieved by subtracting the normalization result from a constant 1. The mean of the eigenvalue difference set is calculated, and the result of performing a negative correlation mapping on the mean of the eigenvalue difference set is recorded as the second eigenvalue. Here, a negative exponential function with a base e is used for mapping. The mean of the first and second eigenvalues is calculated and recorded as the distribution complexity of the target leaf node at the current time. The specific calculation expression for the distribution complexity of the target leaf node at the current time is: Where F is the distribution complexity of the target leaf node at the current time, M is the total number of vectors in the comprehensive vector set, M1 is the first quantity value, M2 is the second quantity value, and S is the total number of data in the feature difference set. Let be the value of the s-th parameter in the set of feature differences, exp() be an exponential function with base e, and M be the value of the parameter. Normalize, In order to Perform negative correlation mapping; This represents the difference in the number of vectors of the two classes among all vectors received by the target leaf node, when A larger value indicates that there are more vectors belonging to the same category in the comprehensive vector set, and the category distribution of the target leaf nodes is less complex; conversely, a smaller value indicates that... The smaller the value, the more similar the number of vectors of the two categories in the composite vector set, and the more complex the category distribution in the target leaf node. This also indicates a higher probability that the target leaf node has not effectively distinguished the samples. In this case, the node should be further split to improve its discriminative ability. When the average mean difference between the two categories in the composite vector set across dimensions or attributes is smaller, that is... The smaller the value, the more blurred the differences in vector attributes and the higher the vector mixing degree in the comprehensive vector set. The nodes should be further split to improve distinguishability; that is, when... smaller and The smaller the value, the higher the complexity of the category distribution in the comprehensive vector set, the more ambiguous the differences in vector attributes in the comprehensive vector set, and the higher the vector hybridity. It also indicates that the distribution complexity of the target leaf node at the current moment is higher, and the target leaf node should be further split to improve the distinguishing ability. smaller and The smaller the value of F, the larger the value of F. Therefore, the larger the value of F, the higher the distribution complexity of the target leaf node at the current moment. This increases the likelihood of the target leaf node triggering a split. In other words, the closer the number of abnormal and normal vectors in the comprehensive vector set and the closer the differences in parameter values under each dimension attribute, the more complex the category distribution, the higher the vector mixing degree, and the more blurred the feature differences in the target leaf node. Therefore, it is more appropriate to continue splitting to improve the distinguishing ability.
[0028] Then, based on the difference in the number of different class vectors in the set formed by the current network performance index value vector and the nearest neighbor vector sequence, the class distribution drift of the target leaf node at the current time is obtained. The specific process for obtaining the class distribution drift of the target leaf node at the current time is as follows: The set formed by the current network performance index value vector and all vectors in the nearest neighbor vector sequence is denoted as the new set. Vectors belonging to the same category in the new set are grouped into the same set and denoted as feature sets. The number of feature sets is at most two. The vectors in one feature set belong to the normal category, and the vectors in the other feature set belong to the abnormal category. The total number of vectors in the feature set with the fewest vectors is denoted as the first feature value, and the total number of vectors in the feature set with the most vectors is denoted as the second feature value. The normalized result of subtracting the second feature value from the first feature value is taken as the category distribution drift of the target leaf node at the current time. The specific expression for calculating the category distribution drift of the target leaf node at the current time is as follows: Where K is the class distribution drift of the target leaf node at the current time, and H is the total number of vectors in the new set. The first eigenvalue is the total number of vectors in the feature set with the fewest vectors. This represents the total number of vectors in the feature set with the largest number of vectors, and is also the second eigenvalue. The range of the ratio to H is -1 to 1. Compared to H, adding 1 gives a range of 0 to 2. Then multiplying by one-half makes the range of K range from 0 to 1. A larger value indicates a significant increase in the number of previously few category vectors, or a noticeable shift in sample distribution within recent nodes. This suggests a decline in the discriminative power of the original segmentation features of the target leaf node, indicating a need to trigger a split to re-divide the feature space and improve the model's generalization ability. The larger K is, the larger K becomes. Therefore, the larger K is, the more it indicates that the number of class vectors that were originally few has increased significantly recently, or that the sample distribution in recent nodes has changed significantly. This means that the discriminative power of the original partitioning features of the target leaf node is decreasing, and it indicates that a split needs to be triggered to repartition the feature space and improve the generalization ability of the model. In other words, the larger K is, the greater the possibility of the target leaf node triggering a split.
[0029] Finally, the mean of the distribution complexity of the target leaf node and the class distribution drift of the target leaf node at the current time are calculated and denoted as the splitting necessity index value of the target leaf node at the current time. The splitting necessity index value of the target leaf node at the current time can reflect whether the target leaf node needs to be split at the current time. The larger the splitting necessity index value, the more it indicates that the target leaf node needs to trigger a split in order to re-divide the feature space and improve the generalization ability and discrimination ability of the model.
[0030] In this embodiment, the specific process of making a splitting decision on the target leaf node at the current moment based on the splitting necessity index value of the target leaf node at the current moment is as follows: it is determined whether the splitting necessity index value of the target leaf node at the current moment is greater than the preset splitting necessity threshold. If so, it is determined that the target leaf node at the current moment needs to trigger a split to improve the algorithm's future discrimination and generalization ability. Otherwise, it is determined that the target leaf node at the current moment does not need to trigger a split. In specific applications, the implementer needs to set the preset splitting necessity threshold according to the value range of the splitting necessity index value, experimental statistics, and other actual conditions. For example, in this embodiment, the preset splitting necessity threshold can be set to 0.7.
[0031] In this embodiment, the process of making a splitting decision on the target leaf node at the current moment based on the Hoeffding inequality is well-known. That is, the process of determining whether the target leaf node at the current moment needs to trigger a split based on the Hoeffding inequality is well-known. After determining whether the target leaf node at the current moment needs to trigger a split, the result of identifying abnormal traffic in the industrial IoT at the current moment can be output. Furthermore, the Hoeffding tree in this embodiment is an incremental decision tree learning algorithm specifically designed for handling data stream classification problems. Compared to existing Hoeffding tree-based abnormal traffic identification processes, the process of identifying abnormal traffic in this embodiment only changes the splitting decision mechanism and does not change other process steps. In other words, this embodiment only improves and optimizes the determination process of whether the target leaf node at the current moment needs to trigger a split, without improving other process steps.
[0032] Thus, this embodiment completes the online identification of abnormal traffic in the Industrial Internet of Things (IIoT). Furthermore, by incorporating a split necessity index value into the split decision-making process, this embodiment can improve the correctness, reliability, and timeliness of the split decision, thereby enhancing the accuracy and reliability of identifying abnormal traffic in the IIoT.
[0033] In summary, this embodiment first obtains the target leaf node. Then, based on the optimal attributes of each vector in the nearest neighbor vector sequence of the target leaf node, and the difference between the optimal and second-best attribute gain of the target leaf node at the current time and the time before the current time, it obtains the confidence characterization value of the target leaf node at the current time. Next, it determines whether the confidence characterization value is greater than a preset confidence threshold. If so, it makes a splitting decision for the target leaf node at the current time according to the Hoeffding inequality. Otherwise, it obtains the splitting necessity index value of the target leaf node at the current time based on the difference in the number of different categories of vectors in the set formed by all vectors reaching the target leaf node, the current network performance index value vector, and the nearest neighbor vector sequence. It then makes a splitting decision for the target leaf node at the current time based on the splitting necessity index value and outputs the identification result of abnormal traffic in the industrial IoT at the current time. Furthermore, by incorporating the splitting necessity index value into the splitting decision process, this embodiment improves the correctness, reliability, and timeliness of the splitting decision, thereby enhancing the accuracy and reliability of identifying abnormal traffic in the industrial IoT.
[0034] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for online identification of abnormal traffic in the Industrial Internet of Things (IIoT), characterized in that, The method includes the following steps: Obtain the target leaf node, which is the leaf node reached by the current network performance index value vector after it is input into the Hofding tree, and the current network performance index value vector is the vector of the Industrial Internet of Things at the current moment; Based on the optimal attributes of each vector in the nearest neighbor vector sequence of the target leaf node, as well as the difference between the optimal attribute gain and the difference between the second-best attribute gain of the target leaf node at the current time and the time before the current time, the confidence characterization value of the target leaf node at the current time is obtained. If the confidence level is greater than a preset confidence threshold, a splitting decision is made for the target leaf node at the current moment according to the Hoeffding inequality. Otherwise, the splitting necessity index value of the target leaf node at the current moment is obtained based on the difference in the number of different categories of vectors in the set formed by all vectors reaching the target leaf node, the current network performance index value vector, and the nearest neighbor vector sequence. A splitting decision is made for the target leaf node at the current moment based on the splitting necessity index value, and the identification result of abnormal traffic in the industrial IoT at the current moment is output.
2. The method for online identification of abnormal traffic in industrial IoT as described in claim 1, characterized in that, The nearest neighbor vector sequence consists of a preset number of network performance index value vectors that are closest in time to the current time to the target leaf node. The optimal attribute of any network performance index value vector is the attribute that can most effectively distinguish the traffic state on the node reached by the network performance index value vector. Any network performance index value vector consists of average packet length, packet rate per second, and protocol distribution ratio.
3. The method for online identification of abnormal traffic in industrial IoT as described in claim 1, characterized in that, The methods for obtaining the confidence representation value of the target leaf node at the current time include: Based on the optimal properties of each vector in the nearest neighbor vector sequence, the flip sequence of the target leaf node at the current time is obtained; Let the time before the current time be t-1. Let the information gain of the optimal attribute of the target leaf node at the current time be minus the information gain of the optimal attribute of the target leaf node at t-1 time, and let the information gain of the suboptimal attribute of the target leaf node at the current time be minus the information gain of the suboptimal attribute of the target leaf node at t-1 time, and let the suboptimal attribute gain be. Based on the flip sequence, the optimal attribute gain difference, and the suboptimal attribute gain difference, the confidence characterization value of the target leaf node at the current time is obtained.
4. The method for online identification of abnormal traffic in industrial IoT as described in claim 3, characterized in that, The method for obtaining the flip sequence of the target leaf node at the current moment includes: If the optimal attribute of the a-th vector in the nearest neighbor vector sequence is the same as the optimal attribute of the (a+1)-th vector, then the value of the a-th parameter in the flipped sequence is assigned to 0; if the optimal attribute of the a-th vector in the nearest neighbor vector sequence is not the same as the optimal attribute of the (a+1)-th vector, then the value of the a-th parameter in the flipped sequence is assigned to 1.
5. The method for online identification of abnormal traffic in industrial IoT as described in claim 3, characterized in that, The method for obtaining the confidence representation value of the target leaf node at the current time based on the flip sequence, the optimal attribute gain difference, and the second-best attribute gain difference includes: The result of negatively correlated mapping of the mean of the flipped sequence is denoted as the first characterization value; Will Let be the optimal attribute gain rate of change, and Let this be denoted as the rate of change of the suboptimal attribute gain. The optimal attribute gain difference. The difference in gain between suboptimal attributes is represented by `max()`, which is the function to find the maximum value. Let be the information gain of the optimal attribute of the target leaf node at time t-1. Let be the information gain of the suboptimal attribute of the target leaf node at time t-1. As a preset constant, the normalized result of the absolute value of the difference between the optimal attribute gain change rate and the second-best attribute gain change rate is recorded as the second characterization value. The weighted sum of the first and second characterization values is used as the confidence characterization value of the target leaf node at the current time.
6. The method for online identification of abnormal traffic in industrial IoT as described in claim 1, characterized in that, The methods for obtaining the splitting necessity index value of the target leaf node at the current moment include: The set of all network performance index value vectors received by the target leaf node up to the current time is denoted as the comprehensive vector set. Vectors of the same type in the comprehensive vector set are divided into the same set and denoted as subsets. The number of subsets is two. Based on the differences in the same attribute parameters among the subsets, a set of feature differences is obtained; Based on the number of vectors in the subset and the feature difference set, the distribution complexity of the target leaf node at the current time is obtained; Based on the difference in the number of vectors of different categories in the set formed by the current network performance index value vector and the nearest neighbor vector sequence, the category distribution drift of the target leaf node at the current time is obtained; The average of the distribution complexity and the category distribution drift is recorded as the splitting necessity index value of the target leaf node at the current time.
7. The method for online identification of abnormal traffic in industrial IoT as described in claim 6, characterized in that, Methods for obtaining the feature difference set include: The mean of all vectors in the subset is denoted as the mean vector of the corresponding subset, and the set of absolute values of the differences of the same attribute parameters between different mean vectors is denoted as the feature difference set.
8. The method for online identification of abnormal traffic in industrial IoT as described in claim 6, characterized in that, The methods for obtaining the distributed complexity of the target leaf node at the current time include: The number of vectors in the subset with the largest number of vectors is denoted as the first quantity value, and the number of vectors in the subset with the smallest number of vectors is denoted as the second quantity value. The result of normalizing the first quantity value minus the second quantity value and then performing negative correlation mapping is denoted as the first feature value. The negative correlation mapping result of the mean of the feature difference set is denoted as the second feature value. The mean of the first feature value and the second feature value is denoted as the distribution complexity of the target leaf node at the current time.
9. The method for online identification of abnormal traffic in industrial IoT as described in claim 6, characterized in that, The class distribution drift of the target leaf node at the current moment is the normalized result of the number of vectors with the fewest classes minus the number of vectors with the most classes in the new set, where the new set is the set formed by the current network performance index value vector and the nearest neighbor vector sequence.
10. The method for online identification of abnormal traffic in industrial IoT as described in claim 1, characterized in that, A method for making a splitting decision on the target leaf node at the current moment based on the splitting necessity index value includes: Determine whether the split necessity index value is greater than the preset split necessity threshold. If it is, determine that the target leaf node needs to trigger a split at the current time. Otherwise, determine that the target leaf node does not need to trigger a split at the current time.
Citation Information
Patent Citations
Private domain traffic peak identification and route scheduling method and system based on time sequence prediction
CN119449630A
Intelligent water affair monitoring management system based on Internet of Things
CN120995025A