Distributed network attack tracing method and system

By clustering and dividing the traffic characteristics between distributed network nodes and building a network topology correlation matrix, combining abnormal traffic transmission indicators and node confidence evaluation values, an attack path priority sequence is generated, which solves the problems of limited abnormal traffic recognition capabilities and lack of node confidence evaluation standards in the existing technology, and achieves accurate positioning of network attack sources and improving network security protection levels.

CN120110708APending Publication Date: 2025-06-06ZHEJIANG BOYAN SAFETY TECHNOLOGY CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510083040.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The ability to identify abnormal traffic characteristics in the prior art is limited, and a systematic traffic analysis framework has been failed to establish, making it difficult to detect and deal with potential threats in a timely manner. The access control mechanism lacks quantitative evaluation standards for node credibility, and it is difficult to dynamically adjust the protection strategy, which reduces the level of network security protection.

Method used

By counting the traffic characteristic values ​​between distributed network nodes, clustering and dividing, and generating node traffic characteristic clustering values; building a traffic transmission relationship diagram between nodes, calculating link weights and correlation degree, and generating a network topology correlation matrix; combining abnormal traffic transmission indicators and node traffic characteristic clustering values, building a node confidence scoring standard, calculating node confidence evaluation values, extracting attack path feature sequences, and generating attack path priority sequences.

Benefits of technology

It has achieved accurate positioning of network attack sources, improved the accuracy of attack behavior detection, established a scientific and reasonable traceability decision-making mechanism, and improved the level of network security protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120110708A_ABST
    Figure CN120110708A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network security, in particular to a distributed network attack tracing method and system, and the method comprises the following steps: carrying out the statistics of flow transmission time delay, bandwidth occupancy rate and data packet quantity among distributed network nodes, carrying out the clustering division of each node flow feature value, and generating a node flow feature clustering value; and on the basis of the node traffic feature clustering value, constructing an inter-node traffic transmission relation graph, and calculating an inter-node link weight and an association degree. According to the method, clustering division is carried out on the flow characteristic values among the distributed network nodes, the flow transmission relation graph among the nodes is constructed, the link weight and the correlation degree are calculated, a complex network structure is converted into a topological correlation matrix capable of being quantitatively analyzed, and basic data support is provided for follow-up attack traceability. Data packet identifiers and timestamp information are extracted based on a time series data analysis method, and abnormal traffic transmission characteristics are identified and statistically analyzed, so that the attack behavior detection accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular to a distributed network attack tracing method and system. Background Art

[0002] Network security is a comprehensive technical field that focuses on protecting computer network systems, data, and infrastructure from various threats and attacks. This field covers multiple technical directions, including network defense, vulnerability detection, security auditing, encryption technology, access control, identity authentication, etc. However, the existing technology has limited ability to identify abnormal traffic characteristics, and has failed to establish a systematic traffic analysis framework, making it difficult to detect and handle potential threats in a timely manner. The access control mechanism lacks quantitative evaluation standards for node credibility, making it difficult to dynamically adjust protection strategies, which reduces the level of network security protection. Summary of the invention

[0003] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a distributed network attack tracing method and system.

[0004] In order to achieve the above object, the present invention adopts the following technical solution, a distributed network attack tracing method, comprising the following steps:

[0005] Count the traffic transmission delay, bandwidth occupancy rate and number of data packets between nodes in the distributed network, cluster the traffic characteristic values ​​of each node, and generate node traffic characteristic clustering values; based on the node traffic characteristic clustering values, construct a traffic transmission relationship diagram between nodes, calculate the link weights and correlation degrees between nodes, and generate a network topology correlation matrix;

[0006] According to the network topology association matrix, the packet forwarding path and traffic transmission sequence between nodes are calculated, the packet identifier and timestamp are extracted, and the traffic transmission timing sequence is generated. Based on the traffic transmission timing sequence, the packet identifier is matched and screened, the abnormal traffic transmission characteristics are counted, and the abnormal traffic transmission index is generated;

[0007] Combining the abnormal traffic transmission index with the node traffic feature clustering value, constructing a node credibility scoring standard, and calculating a node credibility evaluation value; based on the node credibility evaluation value, extracting an attack path feature sequence, and generating an attack path feature index;

[0008] Based on the attack path characteristic index and the node credibility evaluation value, the attack path credibility probability is calculated, the attack paths are prioritized, and an attack path priority sequence is generated.

[0009] Preferably, the steps for obtaining the node traffic feature clustering value are:

[0010] Statistics are collected on the traffic transmission delay, bandwidth utilization and number of packets between nodes in the distributed network. The traffic transmission data of each node is normalized and the weighted average value of the delay, bandwidth utilization variance and absolute deviation of the number of packets between nodes are calculated to generate a basic data table of traffic characteristics.

[0011] According to the traffic feature basic data table, feature merging is performed on the weighted average value of the node's delay, the variance of the bandwidth utilization rate, and the absolute deviation of the number of data packets to form a feature data set;

[0012] According to the characteristic data set, the node traffic characteristic clustering value is calculated, and the formula is:

[0013]

[0014] Among them, C k is the node traffic feature clustering value of the kth cluster, m k is the number of nodes in the kth cluster, t i is the weighted average delay of the ith node, b i is the bandwidth utilization variance of the ith node, d i is the absolute deviation of the number of packets of the ith node, is the mean number of packets in the kth cluster.

[0015] Preferably, the steps of obtaining the network topology association matrix are:

[0016] Based on the node traffic feature clustering value, extract the traffic feature vector between nodes, compare the traffic features between nodes in pairs, calculate the transmission frequency and transmission delay difference between node pairs, and obtain a preliminary node relationship feature data table;

[0017] According to the preliminary node-to-node relationship feature data table, the link weights between nodes are calculated using the following formula:

[0018]

[0019] Among them, W ij is the link weight between node i and node j, F ij is the transmission frequency between node i and node j, T ij is the transmission delay between node i and node j, D i and D j are the traffic feature dimension values ​​of node i and node j respectively, C i and C j are the node traffic feature clustering values ​​of node i and node j respectively;

[0020] Based on the link weight, the nodes are regarded as points in the graph, the link weight is regarded as the weight of the edge, a traffic transmission relationship graph is constructed, and the graph data structure is converted into a matrix to generate a network topology association matrix.

[0021] Preferably, the steps of acquiring the traffic transmission timing sequence are:

[0022] According to the network topology association matrix, the link weights in the matrix are traversed one by one, and the forwarding path and forwarding priority of each link are calculated in combination with the size and directionality information of the link weights to form a forwarding path set between nodes;

[0023] Based on the forwarding path set between the nodes, the path order of the data packet is recorded, the transmission order of the data packet on each path is extracted in turn, and the traffic transmission sequence is generated in combination with the time characteristics of the node link weight;

[0024] Based on the traffic transmission sequence, a data packet identifier and a timestamp are extracted from the data packet information in the transmission path, and the timing information of all node data packets are integrated in chronological order to generate a traffic transmission timing sequence.

[0025] Preferably, the steps of acquiring the abnormal traffic transmission indicator are:

[0026] According to the traffic transmission timing sequence, extract the data packet identifiers and the corresponding timestamps one by one, perform statistics on the distribution of the identifiers, record the occurrence frequency and time interval of each identifier, and obtain an identifier time distribution statistics table;

[0027] Based on the identifier time distribution statistics table, filter the data packet identifiers with abnormal frequency or time interval, combine the timestamps in the traffic transmission timing sequence, extract the time series of abnormal transmission, and form abnormal traffic feature data;

[0028] Based on the abnormal traffic characteristic data, the abnormal traffic transmission index is calculated, and the calculation formula is:

[0029]

[0030] Among them, A is the abnormal traffic transmission index, Q i is the occurrence frequency of the ith packet identifier, CT i is the time interval of the ith packet identifier, E i is the transmission timestamp of the ith packet identifier, is the average value of the transmission timestamps of all abnormal data packets, and k is the total number of abnormal data packet identifiers.

[0031] Preferably, the steps for obtaining the node credibility evaluation value are:

[0032] Combining the abnormal traffic transmission index with the node traffic feature clustering value, statistically analyzing the traffic abnormality performance and traffic feature center value of each node, summarizing the feature deviation of each node, and obtaining a preliminary node abnormal feature data table;

[0033] Based on the preliminary node abnormal feature data table, analyze the correlation between the node abnormal feature and other nodes, select the node set with greater abnormal feature influence, and obtain node comprehensive abnormal score data;

[0034] According to the node comprehensive abnormality score data, the node credibility evaluation value is calculated, and the calculation formula is:

[0035]

[0036] Among them, R i is the credibility evaluation value of the i-th node, A i is the abnormal traffic transmission index of the i-th node, C i is the node traffic feature clustering value of the i-th node, M i and N i are the abnormal feature dimension value and abnormal correlation degree of the i-th node respectively.

[0037] Preferably, the steps of acquiring the attack path characteristic index are:

[0038] Based on the node credibility evaluation value, nodes whose node credibility evaluation value is lower than the threshold are screened, and the traffic transmission links between the nodes below the threshold are analyzed one by one, and the directionality and data transmission status of each link are extracted, and the transmission correlation degree between the nodes is calculated in combination with the link weight, and a set of traffic transmission links between suspicious nodes is generated;

[0039] According to the traffic transmission link set between suspicious nodes, the data transmission paths in the links are extracted one by one, the node sequence and data packet transmission rules on the path are analyzed, and the correlation features between the nodes are sorted and classified to extract the attack path feature sequence;

[0040] Based on the attack path feature sequence, the credibility evaluation value of the node and the transmission link directionality data are called, and the characteristic parameters on each path are counted item by item to generate a feature description of each path to form an attack path feature index.

[0041] Preferably, the steps for acquiring the attack path priority sequence are:

[0042] Based on the attack path characteristic index and the node credibility evaluation value, the node information and characteristic parameters on each attack path are extracted, and weighted combination is performed according to the credibility evaluation value of the nodes in the path, and combined with the abnormal traffic information in the path characteristics, to generate an attack path credibility basic data table;

[0043] According to the attack path credibility basic data table, the credibility probability of each attack path is calculated, and the influence of the node characteristics and abnormal traffic characteristics of the path on the credibility probability is statistically analyzed item by item to obtain the attack path credibility probability analysis result;

[0044] Based on the attack path credibility probability analysis result, the attack paths are prioritized from low to high according to credibility probability, the sequence information of the sorted paths is extracted, and the information is integrated to form an attack path priority sequence.

[0045] The present invention provides a distributed network attack tracing system, comprising:

[0046] The traffic feature analysis module collects the traffic transmission delay, bandwidth occupancy rate and number of data packets between network nodes, and generates node traffic feature clustering values ​​by clustering the data;

[0047] The network relationship building module draws the traffic transmission relationship diagram between nodes based on the node traffic feature clustering value, calculates the link weight and correlation degree, and generates the network topology correlation matrix;

[0048] The traffic sequence analysis module uses the network topology association matrix to track the packet forwarding path and traffic transmission sequence, record the packet identifier and timestamp, and generate the traffic transmission timing sequence;

[0049] The abnormal traffic detection module matches and filters the packet identifiers based on the traffic transmission timing sequence, performs statistical analysis on the traffic transmission characteristics, and generates abnormal traffic transmission indicators;

[0050] The attack path evaluation module combines the abnormal traffic transmission indicators and the node traffic feature clustering values ​​to establish the node credibility scoring standard, calculate the node credibility evaluation value, extract the attack path feature sequence, and generate the attack path priority sequence.

[0051] Compared with the prior art, the advantages and positive effects of the present invention are:

[0052] The present invention clusters the traffic characteristic values ​​between nodes in a distributed network, constructs a traffic transmission relationship diagram between nodes, and calculates link weights and correlation degrees, thereby converting complex network structures into a topological correlation matrix that can be quantified and analyzed, providing basic data support for subsequent attack tracing. Based on the time series data analysis method, the packet identifier and timestamp information are extracted, and the abnormal traffic transmission characteristics are identified and statistically analyzed to improve the accuracy of attack behavior detection. Combined with the node credibility scoring mechanism and the attack path characteristic sequence analysis, a multi-dimensional evaluation system is constructed to achieve accurate positioning of the attack source. By calculating the credibility probability of the attack path and prioritizing it, a scientific and reasonable tracing decision-making mechanism is established to improve the efficiency of tracing. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0055] See also Figure 1 The present invention provides a technical solution, a distributed network attack tracing method, comprising the following steps:

[0056] Statistics are collected on the traffic transmission delay, bandwidth occupancy and number of packets between nodes in the distributed network, and the traffic characteristic values ​​of each node are clustered to generate node traffic characteristic clustering values; based on the node traffic characteristic clustering values, a traffic transmission relationship diagram between nodes is constructed, and the link weights and correlation degrees between nodes are calculated to generate a network topology correlation matrix;

[0057] According to the network topology association matrix, the packet forwarding path and traffic transmission sequence between nodes are calculated, the packet identifier and timestamp are extracted, and the traffic transmission timing sequence is generated. Based on the traffic transmission timing sequence, the packet identifier is matched and screened, the abnormal traffic transmission characteristics are counted, and the abnormal traffic transmission index is generated;

[0058] Combine the abnormal traffic transmission index and the node traffic feature clustering value to build a node credibility scoring standard and calculate the node credibility evaluation value; based on the node credibility evaluation value, extract the attack path feature sequence and generate the attack path feature index;

[0059] Based on the attack path characteristic indicators and node credibility evaluation values, the credibility probability of the attack path is calculated, the attack paths are prioritized, and the attack path priority sequence is generated.

[0060] The steps to obtain the node traffic feature clustering value are as follows:

[0061] Statistics are collected on the traffic transmission delay, bandwidth utilization and number of packets between nodes in the distributed network. The traffic transmission data of each node is normalized and the weighted average value of the delay, bandwidth utilization variance and absolute deviation of the number of packets between nodes are calculated to generate a basic data table of traffic characteristics.

[0062] According to the basic data table of traffic characteristics, the weighted average value of node delay, variance of bandwidth utilization and absolute deviation of number of data packets are merged to form a characteristic data set;

[0063] According to the characteristic data set, the node traffic characteristic clustering value is calculated. The formula is:

[0064]

[0065] Among them, C k is the node traffic feature clustering value of the kth cluster, m k is the number of nodes in the kth cluster, t i is the weighted average delay of the ith node, b i is the bandwidth utilization variance of the ith node, d i is the absolute deviation of the number of packets of the ith node, is the mean number of packets in the kth cluster.

[0066] Specifically, the basic data involved in the statistics of traffic transmission delay, bandwidth occupancy and number of data packets between distributed network nodes are mainly obtained from the monitoring process deployed in each node area. Each node area collects and records bandwidth occupancy and delay data within a preset time interval. Combined with the monitoring results of bandwidth occupancy, it can be determined that its variation range is usually between 0% and 90%. For delay information, the actual delay value in the range of 0ms to 300ms is measured by sending test data between multiple nodes in a directional manner. Then, the recorded delay data and bandwidth occupancy data are compared respectively, and abnormal values ​​exceeding the specified abnormal range are screened out (the abnormal range here is obtained by comparing the peak and extreme value quantiles of the historical distribution by generating the historical distribution of delay and bandwidth occupancy through continuous monitoring for one month. For example, the delay data is segmented into histograms and intercepted at the 95% position of the quantile. The value exceeding the interception value is an abnormal point, and the same is true for bandwidth occupancy). Then, the number of data packets in each node in a recent period of time is recorded at a stable sampling period, and its absolute value is compared with the value of the bandwidth occupancy rate. A threshold comparison is implemented within the data range of 0 to thousands (here the threshold is set by the variance range of the mean number of data packets collected in the previous month plus or minus a multiple, and the variance range is obtained by cumulative statistics). After obtaining the transmission records of all nodes, the traffic transmission data of each node is normalized one by one. The method is to first divide the bandwidth occupancy rate, delay and number of data packets by the difference between their respective maximum values ​​and minimum values, so that the data is mapped to the normalized interval between 0 and 1. After obtaining the normalized result, the weighted average value of delay, the variance of bandwidth utilization and the absolute deviation of the number of data packets are further calculated for each node. These parameters are obtained based on the normalization results of the previous step combined with the fluctuations of all nodes in the period for weighted calculation (for example, in the weighted average value processing of delay, a low weight can be given to high-speed nodes, and a high weight can be given to nodes with large delay fluctuations. The weight is determined by continuously observing the standard deviation of node delay fluctuations and normalizing them). Finally, these values ​​are summarized to generate a basic data table of traffic characteristics.

[0067] According to the basic data table of traffic characteristics obtained in the previous step, feature merging is performed for the weighted average value of delay, bandwidth utilization variance and absolute deviation of the number of data packets for each node. The specific execution process includes reading the weighted average value of delay of each node recorded in the table in turn and comparing it with the corresponding value of the adjacent node. By counting the difference of the weighted average value of delay between adjacent nodes and combining the discrete degree of bandwidth utilization variance, it is judged whether there is abnormal numerical distribution. For bandwidth utilization variance, it is checked item by item based on the quantized variance array in the record, and the node data with a variance higher than the specified standard is screened out. The specified variance standard comes from the distribution law obtained by continuous monitoring of bandwidth utilization for many days under daily business conditions. The specific method is to establish a variance between 0 and 10. The variance scale range is divided into several intervals, the frequency statistics of the bandwidth utilization variance in each interval are calculated, and the interval with the highest frequency of occurrence is taken as the main interval. Data beyond this range is regarded as data that needs special attention. Then the screening results are combined with the absolute deviation of the number of data packets of each node. The absolute deviation of the number of data packets is compared longitudinally with the transmission records of the corresponding nodes in the same time period. By superimposing multiple records, it is determined whether there is a situation that significantly exceeds the historical average deviation of the same node, and the node records with abnormal deviations are marked. After the sorting is completed, the parameter set after feature merging processing can be obtained. These parameters are stored in the same data structure to provide a unified input for the calculation of subsequent feature data sets, and finally form a feature data set.

[0068] The benefit of the formula is that by simultaneously considering the ratio of the weighted average of delay to the variance of bandwidth utilization, as well as the absolute deviation of the node at the level of data packet quantity, it can comprehensively measure the multiple differences of each node in the process of network traffic transmission, and further reflect the overall distribution of nodes within the kth cluster in multi-dimensional characteristics, thereby providing more accurate input parameters for subsequent node credibility evaluation.

[0069] In the formula, t i It represents the weighted average delay of the ith node, in ms, and the value range is usually between 10ms and 300ms. In the current example, t is determined based on the basic data table of traffic characteristics obtained above. 1 =20, t 2 =25, t 3 =15, t 4 =30. These values ​​are obtained by recording the round-trip delays between a large number of nodes during a continuous monitoring period and performing weighted processing. The weights of the weighted processing come from the distribution comparison of the node's real-time traffic and the historical average traffic. i It represents the bandwidth utilization variance of the ith node. The value range is generally between 0.01 and 0.1. In the current example, b is set according to the bandwidth utilization variance collection result obtained previously. 1 =0.02, b2 =0.03, b 3 =0.018, b 4 =0.025, where each variance value is derived from the discreteness statistics of the continuous bandwidth monitoring data for multiple days, and is obtained by dividing the bandwidth utilization fluctuation range into segments and then calculating the variance. i is the absolute deviation of the number of packets of the ith node, ranging from 10 to 200. In the current example, it is denoted as d based on the basic data of traffic characteristics obtained previously. 1 =50,d 2 =60, d 3 =45,d 4 =70, which is calculated by recording the fluctuation of the number of packets per minute of the node within a certain time window and then comparing it with the historical average value of each node within the same time range. is the mean number of packets in the kth cluster, which can be obtained from the d i The values ​​are accumulated and averaged. In the current example, the kth cluster contains four nodes, so The specific value is obtained by summing the absolute deviation of the number of data packets of all nodes in the cluster and dividing it by the total number of nodes. k | is the number of nodes in the kth cluster. The cluster in the current example contains four nodes, so |m k |=4.

[0070] Calculation process: First, sum each part of the term: The first node: In calculation hour,

[0071]

[0072] Then calculate

[0073]

[0074] The sum of the two items is 20000+0.111=20000.111.

[0075] Second node:

[0076]

[0077]

[0078] Add them together and you get 20833.333+0.0667≈20833.3997;

[0079] The third node:

[0080]

[0081]

[0082] The sum of the two items is 12500+0.2=12500.2;

[0083] 4th node:

[0084]

[0085]

[0086] Add together to get 36000+0.2444=36000.2444;

[0087] Add the results of the above four nodes: 20000.111+20833.3997+12500.2+

[0088] 36000.2444≈89333.9551;

[0089] Then divide by |m k |=4,

[0090]

[0091] Finally, we calculate the square root:

[0092]

[0093] The result shows that the node traffic feature clustering value is approximately 149.4449 under the current clustering. The larger the value, the higher the comprehensive ratio of delay and bandwidth utilization variance. At the same time, there is also a certain increase in the deviation of the number of data packets. For the subsequent node credibility evaluation step, this result can be used to distinguish the overall transmission stability of the cluster in the network environment.

[0094] The steps to obtain the network topology association matrix are:

[0095] Based on the clustering value of node traffic characteristics, the traffic characteristic vector between nodes is extracted, the traffic characteristics between nodes are compared pairwise, the transmission frequency and transmission delay difference between node pairs are calculated, and the preliminary node relationship characteristic data table is obtained;

[0096] According to the preliminary node-to-node relationship feature data table, the link weights between nodes are calculated using the following formula:

[0097]

[0098] Among them, Wij is the link weight between node i and node j, F ij is the transmission frequency between node i and node j, T ij is the transmission delay between node i and node j, D i and D j are the traffic feature dimension values ​​of node i and node j respectively, C i and C j are the node traffic feature clustering values ​​of node i and node j respectively;

[0099] Based on the link weight, the nodes are regarded as points in the graph, the link weight is regarded as the weight of the edge, a traffic transmission relationship graph is constructed, and the graph data structure is converted into a matrix to generate a network topology association matrix.

[0100] Specifically, based on the node traffic feature clustering value, referring to the clustering category of each node in the clustering result generated in the previous step, by reading the node number and the corresponding traffic feature set recorded in the clustering result, the traffic feature vectors between the nodes are extracted one by one in the comparison order, and then these node traffic feature vectors are compared one by one according to the dimensional position during the comparison, and the feature value differences including bandwidth utilization, number of data packet transmissions, and round-trip delay are recorded, and the difference information is compared with the pre-established effective range. For example, when comparing bandwidth utilization, it is compared with the usage interval of 0% to 90%. If it is observed that the bandwidth utilization deviates from the recorded , it is marked in the current node comparison row, and then the delay value is compared similarly, the delay value is compared with the interval of 0ms to 300ms, and the node combination whose delay exceeds the specified range is recorded. At the same time, the quantity difference is counted according to the increase or decrease in the number of data packets transmitted between adjacent nodes. The characteristic value difference between each node pair is continuously accumulated and recorded, and compared with the quantile threshold in the aforementioned feature set. The quantile threshold is formed by sorting the characteristic value distribution obtained by monitoring the same type of nodes over a period of time and then selecting the percentile point. Finally, after comparing all nodes pairwise, a preliminary node relationship feature data table is formed.

[0101] The benefit of the formula is that it takes into account the differences in transmission frequency and transmission delay between nodes, and also combines the absolute difference between the node traffic feature dimension value and the node traffic feature cluster value, thereby taking multiple factors into account when measuring the link weight between nodes.

[0102] In the formula, F ij It represents the transmission frequency between node i and node j. The unit can be set to the number of data transmissions that occur per minute, usually ranging from 10 to 1000 times. This example uses the preliminary node relationship feature data table obtained above to count the transmission frequencies between node i and node j in multiple time periods within a day and take the average value, which is recorded as F 12=400, F 13 =120, F 23 =350 etc. ij represents the transmission delay between node i and node j, in ms, and the monitoring range is between 10ms and 300ms. This example extracts multiple round-trip delay values ​​between nodes from the same data table, and obtains T by weighted averaging the discrete delay data of a monitoring period. 12 =60, T 13 =50, T 23 =70 etc. i and D j are the traffic feature dimension values ​​of node i and node j, respectively. The values ​​are usually between 0.5 and 5.0, which are used to quantify the comprehensiveness of the feature dimension of the node in terms of bandwidth, delay, and data packet. In this example, after combining the feature parameters obtained in the previous stage and superimposing the node historical dimension data, D is selected. 1 =2.0, D 2 =2.5, D 3 =3.0. C i and C j are the node traffic feature clustering values ​​of node i and node j, which are usually in the range of 10 to 200. In this example, C is extracted from the node traffic feature clustering results obtained previously. 1 =50, C 2 =100, C 3 =60, etc.

[0103] Calculation process: Select the parameters of node 1 and node 2 for calculation, let F 12 =400, T 12 =60, D 1 =2.0, D 2 =2.5, C 1 =50, C 2 =100, then:

[0104]

[0105] |C 1 -C 2 |=|50-100|=50

[0106] ln(1+|C 1 -C 2 |)=ln(1+50)=ln(51)≈3.9318

[0107] After adding, we get:

[0108] W 12 =152.142+3.9318≈156.0738

[0109] Then perform the same formula operation on nodes 1 and 3, and nodes 2 and 3, and we can get W respectively. 13 With W 23 And incorporate the results into the overall link weight summary table.

[0110] The results show that the higher the value, the greater the frequency and delay differences between nodes, and the corresponding clustering value differences are more significant. For example, when it is greater than 150, it often means that the nodes have a large gap in transmission level or feature clustering value. When it is less than 100, it means that the nodes are closer. The above link weights can be used to further construct the network topology association matrix.

[0111] Based on the value represented by the link weight, referring to the previously obtained information such as the node pair transmission frequency, bandwidth utilization, and clustering value, the weight distribution of the node pair is used as a connection reference. The weight values ​​marked in the node pair list are extracted one by one and converted into the corresponding graph edge length or connection strength parameter. Then, the coordinate position of each node is first constructed in a visual layout, the node sequence number and traffic characteristic attributes are marked, and then the node pair weight data table is traversed and the weight value range is read item by item. For example, the node pairs with a weight greater than 100 are divided into a strong connection interval, while the node pairs with a low weight are divided into a strong connection interval. The node pairs with a delay greater than 50 are classified into the weaker connection interval. Each group of node pairs is judged in turn. When the corresponding interval is met, the edge is drawn and the value range is marked on the edge. At the same time, the situation where the transmission delay between nodes is in the range of 10ms to 300ms is noted to distinguish the meaning of connections at different levels. Finally, the formed node and edge information are counted into a matrix structure, and the row and column coordinates corresponding to each unit position in the matrix are mapped one by one to the node number, the weight value is written into the matrix unit, and the unit position with empty connection is set to zero or left blank, thereby obtaining the network topology association matrix.

[0112] The steps for obtaining the traffic transmission timing sequence are as follows:

[0113] According to the network topology association matrix, the link weights in the matrix are traversed one by one, and the forwarding path and forwarding priority of each link are calculated by combining the size and directionality information of the link weight to form a forwarding path set between nodes;

[0114] Based on the forwarding path set between nodes, the path order of the data packet is recorded, the transmission order of the data packet on each path is extracted in turn, and the traffic transmission sequence is generated by combining the time characteristics of the node link weight;

[0115] Based on the traffic transmission sequence, the packet identifier and timestamp are extracted from the packet information in the transmission path, and the timing information of all node packets is integrated in chronological order to generate the traffic transmission timing sequence.

[0116] Specifically, according to the network topology association matrix obtained above, the link weight values ​​stored therein are traversed one by one, and the starting node and target node numbers of each link are first extracted from the recorded link directionality information, and then sorted according to the weight value of each link. When sorting, the distribution range of the link weight between 0 and 1000 is compared. If the value is in a higher section for the same node connection, it will be marked as a high priority in subsequent statistics. If it is lower than the result of subtracting a variance multiple from the mean value obtained by multiple monitoring, it is considered that the link is in a low priority state in this stage. The variance multiple here is obtained by counting the traffic fluctuation degree of the same network in different time periods. Then, combined with the directionality between nodes, the starting node is connected and compared according to its possible predecessor link number, and a recursive formula is executed for each link. The labeling process records the complete path number sequence from the earliest node to the target node. In this process, if the link weight exceeds a threshold upper limit based on the historical peak value (the threshold is obtained by selecting the 95% quantile from the monitoring data in the past three months), it is marked as an extreme link in the path number sequence. If the link weight is lower than the threshold but higher than the mean, it is marked as a median link. After the labeling is completed, the node numbers of each path are further hierarchically summarized according to the sorting results of the links in the path. Each path has a list of feasible forwarding nodes. Combined with the priority marked in the previous step, the high-priority link is placed in front for recording, and the median link and the low-priority link are accumulated in turn. Finally, the forwarding path order corresponding to each link is obtained and recorded in a set to form a forwarding path set between nodes.

[0117] Based on the forwarding path set between nodes, we first collect the node numbers contained in each path and correspond them to the path sequence. Then, we classify the nodes under the path into the transmission sequence index table in sequence to store the adjacent transmission relationship between nodes. Then, we combine the time characteristics of the fluctuation interval of the weight value between 0 and 500 to count the time characteristics. Each link has different time series records in the sampling period. These time series can be matched with the previously obtained path numbers. If the weight value is detected to be one variance multiple higher than the mean in a certain time period, it indicates that the path may have a phenomenon of concentrated data transmission in the time period, which needs to be marked at adjacent times. If such a value appears for multiple consecutive monitoring periods, these time periods can be divided out in the index table. When summarizing, the data packets in the corresponding time period are arranged in order to identify the transmission order of the path at different times. Finally, the time characteristics of all paths are spliced ​​and connected in series in time order to obtain a continuous dynamic transmission change sequence between nodes, which is recorded as a traffic transmission sequence.

[0118] Based on the traffic transmission sequence, the relevant data packet information of each node is retrieved from the transmission path, including the node number marked in the sequence and the transmission order between the adjacent node numbers. The entry and exit time of the data packet at each node is determined according to the monitoring records, and the entry and exit time is extracted in the form of timestamps. Then all data packet identifiers and timestamps are summarized at minute-level or finer intervals. For high concurrency scenarios, the recording granularity of 10ms can be combined. If the data packet identifier jumps or exceeds the valid range within the adjacent timestamp interval, it can be marked with reference to the abnormal quantile value extracted according to the daily operation period. The quantile value is configured to count the 90% position of the transmission volume distribution of the node during normal periods. When this value is exceeded, it is recorded until the timestamps and identifiers of all nodes are arranged in succession. Finally, they are sorted in the index table in order from early to late, and the timestamps and identifiers are merged into a timing information string and compared with the path information between nodes. After completing the timing splicing, the complete traffic transmission timing sequence is obtained.

[0119] The steps to obtain abnormal traffic transmission indicators are as follows:

[0120] According to the traffic transmission timing sequence, extract the data packet identifier and the corresponding timestamp one by one, make statistics on the distribution of the identifier, record the occurrence frequency and time interval of each identifier, and obtain the identifier time distribution statistics table;

[0121] Based on the identifier time distribution statistics table, filter the packet identifiers with abnormal frequency or time interval, combine the timestamps in the traffic transmission timing sequence, extract the time series of abnormal transmission, and form abnormal traffic feature data;

[0122] Based on the abnormal traffic characteristic data, the abnormal traffic transmission index is calculated. The calculation formula is:

[0123]

[0124] Among them, A is the abnormal traffic transmission index, Q i is the occurrence frequency of the ith packet identifier, CT i is the time interval of the ith packet identifier, E i is the transmission timestamp of the ith packet identifier, is the average value of the transmission timestamps of all abnormal data packets, and k is the total number of abnormal data packet identifiers.

[0125] Specifically, according to the traffic transmission timing sequence obtained above, the data packet identifiers and timestamps are extracted one by one, and the number of occurrences and the interval between two adjacent occurrences of different identifiers are recorded respectively. Then, these numbers are compared with the interval values. When analyzing, the number of occurrences from 0 to 500 and the interval range from 1ms to 1000ms are compared. If the number of occurrences of an identifier is higher than the quantile value selected from the aforementioned monitoring history, it is regarded as a high-frequency identifier and specially marked. The quantile value is determined by summarizing the number of occurrences of similar nodes in 30 days of continuous monitoring and taking about 95% of the positions. At the same time, if the interval value of an identifier is higher than the quantile value selected from the aforementioned monitoring history, it is regarded as a high-frequency identifier and specially marked. If the fluctuation exceeds a reference interval set by experience, it is judged as an abnormal interval. The specific upper and lower limits of the reference interval are obtained based on long-term observations of network traffic load and abnormal distribution analysis. After the statistics are completed, the occurrence frequency and time interval values ​​corresponding to each identifier are stored in the same record structure. During the comparison process, the maximum fluctuation range of multiple time periods will be compared and whether the identifier appears intensively in a short period of time will be reviewed again. If the appearance time is too concentrated, the proportion will be additionally noted in the record and compared horizontally with other identifiers. When all identifiers are processed, they are arranged in the order of identifiers or the order of appearance, and the identifier time distribution statistics table is summarized and generated.

[0126] Based on the identifier time distribution statistics table, when screening out packet identifiers with abnormal frequency or time interval, first compare the record row where the identifier is located with the statistical data of each time segment to observe whether the identifier shows a peak or obvious increase or decrease in different time windows. If the packet frequency is higher than the median value extracted from the previous node record plus twice the variance range, it is marked as a high-frequency anomaly. If the time interval fluctuation exceeds the reference interval, it is marked as an interval anomaly. The reference interval is obtained by defining the most common fluctuation range of the node packet transmission interval within a multi-day monitoring cycle and calculating the quantile position therein for 30 consecutive days. Subsequently, these marked packet identifiers are paired with the corresponding timestamps to form a time series, which is compared with the previously recorded time series information. The packet identifiers containing abnormal intervals or abnormal occurrence frequencies are concentrated and arranged on the same time axis. If abnormal records of the same identifier appear in several consecutive monitoring cycles, they are placed in the abnormal summary list for subsequent combined analysis. Finally, the time series content of these abnormal identifiers is associated and summarized to form abnormal traffic feature data.

[0127] The benefit of the formula is that by combining the frequency and time interval of each data packet identifier and its degree of deviation on the time axis, it can provide more targeted measurement results for determining network traffic anomalies.

[0128] The steps to obtain each parameter are as follows: Q iIndicates the frequency of occurrence of the ith data packet identifier, and the numerical range is usually between 10 and 500 times. It is obtained by dividing the total number of times the same identifier appears in several monitoring cycles by the number of monitoring cycles through the identifier time distribution statistics table generated previously; CT i It represents the time interval of the ith data packet identifier, in ms, generally in the range of 10ms to 2000ms. The time difference between two consecutive occurrences of the identifier is recorded and the average value is taken. The degree of fluctuation can be further evaluated in combination with the variance. i The transmission timestamp of the ith data packet identifier is recorded in the specific time point extracted during the daily monitoring cycle, usually in the format of yyyy-MM-ddHH:mm:ss.sss, and then converted into relative time or absolute seconds for calculation; It is the average value of the transmission timestamps of all abnormal data packets, obtained by accumulating the timestamp values ​​of all abnormal identifiers and then dividing by the number of data packet identifiers. The unit is the same as E i The same; k is the total number of abnormal data packet identifiers, which is obtained by counting the identifiers that screen out the abnormal frequency of occurrence and the abnormal time interval, and usually ranges from 1 to more than a hundred.

[0129] Calculation process: First list the example values ​​of each parameter: let k = 4, Q 1 =80, Q 2 =100, Q 3 =50, Q 4 =120, CT 1 =400ms, CT 2 =300ms, CT 3 =500ms, CT 4 = 250ms, convert the timestamp into a second-level integer and let E 1 =19025, E 2 =19120, E 3 =19200, E 4 =19085, accumulated E 1 +E 2 +E 3 +E 4 =19025+19120+19200+19085=76430, / 4 = 19107.5 seconds; calculation Example: |19025-19107.5|=82.5, 9.0885; calculated for all i separately: For example CT 1 =400, Add We get 25.0885, and the other items are summarized in the same way; sum up the results of all i and divide by k = 4 to get A. The complete multi-level operation process can be abbreviated as:

[0130]

[0131] Finally, A≈several values ​​are obtained to characterize the overall level of abnormal traffic transmission indicators within the monitoring range.

[0132] The result shows that the higher the value, the more prominent the deviation between the frequency and time interval. If it exceeds a certain quantile, it can be determined that the batch of identifiers has obvious anomalies. If it is lower than the mean, it means that the overall situation is relatively controllable. By combining A with other indicators in the future, different degrees of anomalies can be more comprehensively distinguished.

[0133] The steps to obtain the node credibility evaluation value are as follows:

[0134] Combine the abnormal traffic transmission index with the node traffic feature clustering value, perform statistical analysis on the traffic anomaly performance and traffic feature center value of each node, summarize the feature deviation of each node, and obtain a preliminary node abnormal feature data table;

[0135] Based on the preliminary node abnormal feature data table, analyze the correlation between the node abnormal features and other nodes, select the node set with greater abnormal feature influence, and obtain the node comprehensive abnormal score data;

[0136] According to the node comprehensive abnormality score data, the node credibility evaluation value is calculated, and the calculation formula is:

[0137]

[0138] Among them, R i is the credibility evaluation value of the i-th node, A i is the abnormal traffic transmission index of the i-th node, C i is the node traffic feature clustering value of the i-th node, M i and N i are the abnormal feature dimension value and abnormal correlation degree of the i-th node respectively.

[0139] Specifically, combining the abnormal traffic transmission index obtained previously with the node traffic feature clustering value, read the basic feature data such as bandwidth utilization, delay fluctuation and data packet transmission frequency of each node in the specified monitoring period, and identify whether there is a feature deviation that does not conform to the preset range by comparing the records of each node in the same period. If the bandwidth or delay value of a node in multiple monitoring periods deviates from the center value by more than a threshold value obtained from empirical statistics, it will be marked in the data of the corresponding node. The threshold used in empirical statistics is derived from the quantile analysis of the high-frequency fluctuation range during the three-month continuous monitoring period and the most frequently occurring range is selected as the main reference, and then a horizontal comparison is made on the increase or decrease in the frequency of data packet transmission. For example, the frequency range of a node is set to 0 to 800 times per minute. If it is observed that the transmission frequency of a node in a short period of time is higher than the upper limit of the range or fluctuates violently and lasts for multiple time periods, the characteristic performance of the node is included in the abnormal deviation list. Subsequently, the information of all nodes included in the list is summarized and compared with the overall node data information to find out the nodes that deviate in multiple aspects. If a node has obvious deviations in the three elements of bandwidth, latency, and data packet frequency, it is marked as a high priority node in the system. If it deviates only in a single element, it is marked as a medium priority node. Finally, the deviations of these nodes are arranged in a unified format and form a data structure to obtain a preliminary node abnormal feature data table.

[0140] Based on the preliminary node abnormal feature data table, the abnormal features of each node are compared with the feature performance of other nodes according to the node number. First, the abnormal feature impact is stratified. If a node shows a continuous increase in bandwidth accompanied by an abnormal increase in latency, it can be considered in the analysis that it has a greater impact on the overall network transmission. By comparing it with the bandwidth and latency distribution of other nodes one by one, it is determined whether there are multiple overlapping abnormal time periods or data packet surges. If multiple overlaps are confirmed, a higher abnormal correlation value is given to the node in the process of node comprehensive abnormality scoring. When stratifying, the statistical results obtained in multiple monitoring cycles will be combined. The upper and lower limits of the values ​​are used for judgment. For example, the reference bandwidth range is 0% to 90%, and the delay range is 0ms to 300ms. The nodes above the 80% percentile or exceeding 250ms are regarded as potential large impact points. Then, their performance in data packet sending interval or occurrence frequency is compared in turn. If the occurrence frequency is also higher than the percentile interval obtained from historical monitoring, the abnormal level is cumulatively increased. After sorting out such situations, the node set with significant abnormal feature impact can be screened out, and the score values ​​of each node in this set in different feature dimensions are weighted or summarized. After such a screening and statistical comparison, the node comprehensive abnormal score data is obtained.

[0141] The benefit of the formula is that it considers the relative proportion of the node's abnormal traffic transmission index and the node traffic feature clustering value at the same time, and incorporates the abnormal feature dimension value and the abnormal correlation degree into the calculation in the form of Euclidean distance, thereby reflecting the credibility of the node in a unified probability output.

[0142] In the formula, M i Indicates the abnormal feature dimension value of the i-th node, usually ranging from 0.5 to 5.0. It is determined by comparing the comprehensive deviation degree of the node in bandwidth, latency, and data packet fluctuation. The larger the value, the greater the deviation. i It represents the abnormal correlation of the ith node, which is generally between 0.1 and 2.0. It is used to quantify the degree of co-occurrence of abnormal features between this node and other nodes. It is obtained by comparing the probability or frequency of simultaneous occurrence of abnormalities between nodes.

[0143] Calculation process: Set the parameter of the first node to A 1 =50, C 1 =100, M 1 =2.0, N 1 =1.0, calculate first Then Add the two together to get 0.5+2.236=2.736, then execute exp(-2.736)≈0.0647, and substitute this value into the formula:

[0144]

[0145] If the parameter of the second node is set to A 2 =30, C 2 =50, M 2 =3.0, N 2 =1.5, the credibility evaluation value can be obtained and compared by following the same steps.

[0146] The results show that when A high proportion or When the value is large, the exponential term will tend to grow in a negative direction, thereby reducing the node credibility assessment value. When the calculated result is greater than 0.9, it often means that the node is in a relatively trustworthy state in the current network. When it is less than 0.5, it indicates that its abnormal characteristics are more significant and need to be paid special attention in subsequent links.

[0147] The steps to obtain the attack path characteristic indicators are as follows:

[0148] Based on the node credibility evaluation value, the nodes with credibility evaluation values ​​lower than the threshold are screened, and the traffic transmission links between the nodes below the threshold are analyzed one by one, the directionality and data transmission status of each link are extracted, and the transmission correlation degree between the nodes is calculated in combination with the link weight, and the traffic transmission link set between the suspicious nodes is generated;

[0149] According to the traffic transmission link set between suspicious nodes, the data transmission paths in the links are extracted one by one, the node sequence and data packet transmission rules on the path are analyzed, and the correlation features between the nodes are sorted and classified to extract the attack path feature sequence;

[0150] Based on the attack path feature sequence, the credibility evaluation value of the call node and the transmission link directionality data are counted item by item on each path to generate a feature description of each path and form an attack path feature index.

[0151] Specifically, based on the node credibility evaluation value obtained above, the evaluation results of all nodes are compared one by one, and the nodes whose credibility evaluation values ​​are lower than the threshold are locked. The threshold is derived from the average credibility quantile analysis results accumulated during the long-term monitoring of the system. Usually, the low quantile position of about 10% is selected as the dividing point. If the evaluation result of a node is lower than this quantile value for many times in multiple observations, it is marked as a suspected high-risk node in the subsequent analysis. Then, combined with the historical traffic transmission records between low-credibility nodes, the transmission directionality between nodes is statistically analyzed. The directionality indicates the data flow direction between nodes. If a node is found to be multiple If a node sends a large amount of data to another node for the second time, and after checking, its time distribution and bandwidth occupancy are both in the high-load range, this node pair is marked as an associated combination that needs attention, and then the link weights of these associated combinations are compared with the actual transmission conditions in adjacent time periods. The link weights are usually in the numerical range of 10 to 500. If the link weight exceeds a high value established through experience (the high value is calculated based on the quantiles of large-volume transmission in the past six months), this link is also included in the suspicious list. Finally, the directional data of the marked nodes and their associated links are recorded in turn, and summarized to form a set of traffic transmission links between suspicious nodes.

[0152] According to the set of traffic transmission links between suspicious nodes, the data transmission path information in each link is read. The start node and target node numbers are first extracted and arranged in order. The intermediate nodes between each number are also recorded synchronously to form a complete node sequence. Then, the fluctuation of data packet transmission is counted in each node sequence. For example, the difference between the bandwidth occupancy rate and the number of data packets between nodes in adjacent monitoring cycles is counted, and these differences are compared with the bandwidth range of 0% to 90% or the range of 0 to thousands of data packets. If the difference is found to be greater than the peak boundary based on the historical distribution (this peak boundary comes from the quantile extraction of multi-day monitoring data), the node association is marked. In addition, the delay in the node sequence is checked to see if there are multiple high delay records close to 300ms. If so, these delay records are linked to the data packet rules of the corresponding nodes for comprehensive judgment. At the same time, the transmission association features between nodes are classified and sorted out, so as to identify the feature combination related to the attack behavior in subsequent operations. Finally, the sorted node sequence is merged with the data packet transmission rule to obtain the attack path feature sequence.

[0153] Based on the attack path feature sequence, the node credibility evaluation values ​​and the directional data of the transmission link obtained in the previous step are retrieved in turn, and the bandwidth occupancy, delay record, and data packet density on each path are arranged in chronological order. Each feature corresponds to the relative position of the current node in the path and the upstream and downstream node numbers. If the credibility evaluation value is lower than the low percentile established in the previous article, a reminder is added to the list. If the directionality indicates that a node continuously outputs a large amount of data to other low-credibility nodes, it is also specially marked. Then, the abnormal conditions of the entire path are classified and counted based on these marking information. Finally, the bandwidth, delay, data packet density, and credibility of each path are combined into a multi-dimensional feature description to form an attack path feature indicator.

[0154] The steps for obtaining the attack path priority sequence are:

[0155] Based on the attack path characteristic indicators and node credibility evaluation values, the node information and characteristic parameters on each attack path are extracted, and weighted combination is performed according to the credibility evaluation values ​​of the nodes in the path, and combined with the abnormal traffic information in the path characteristics to generate the attack path credibility basic data table;

[0156] According to the attack path credibility basic data table, the credibility probability of each attack path is calculated, and the impact of the node characteristics and abnormal traffic characteristics of the path on the credibility probability is statistically analyzed item by item to obtain the attack path credibility probability analysis result;

[0157] Based on the attack path credibility probability analysis results, the attack paths are prioritized from low to high according to the credibility probability, the sorted path sequence information is extracted, and integrated to form the attack path priority sequence.

[0158] Specifically, based on the attack path characteristic indicators and node credibility evaluation values ​​obtained previously, the node numbers and corresponding credibility evaluation value ranges in each attack path are first read, and these nodes are verified according to the credibility interval from 0 to 1. The nodes that are not in this interval or are continuously lower than the critical threshold obtained by the quantile analysis are marked, and then the information of each node in the path is extracted in turn, including the node identification number and the abnormal traffic record associated with it. The number of data packets, delay and bandwidth occupancy reflected in the corresponding abnormal traffic record are compared with the range of 0 to thousands of data packets, the delay range of 10ms to 300ms and the bandwidth utilization rate of 0% to 90%. Additional annotations are set for the parts that exceed the high quantile threshold selected after several days of monitoring. Next, combined with the node in the path, the The order of nodes in the path and the credibility evaluation value are combined in a weighted manner. The setting of the weight depends on the relative position of the node in the attack path and the abnormal frequency of the node in multiple monitoring cycles. For example, if the node touches the top of the quantile obtained by the previous statistics multiple times in bandwidth occupancy, it will be increased in the corresponding weighted process. Similarly, the weight accumulation result of each path at the node level is recorded, and the abnormal traffic information is compared item by item. For example, it is determined whether the aforementioned record exceeds the acceptable packet number fluctuation range or the bandwidth high occupancy range during the monitoring period. When a path connecting multiple such highly abnormal nodes is observed, this path is marked as a key focus item, and then all the marks and weight accumulation calculation results are summarized and arranged into a data structure table to form an attack path credibility basic data table.

[0159] According to the attack path credibility basic data table, the credibility and abnormal traffic characteristics are listed for each node in the attack path. Then, the number of abnormal or suspicious states of the node is obtained from the previously recorded time periods. These numbers are compared with the preset range of 0 to 10. If the frequency is higher than the reference threshold selected based on the historical data quantile, it is recorded as a high-frequency abnormal node. If the abnormality only occurs in individual cycles, it is recorded as a low-frequency abnormal node. Then, the credibility evaluation value of each node in the path is combined with the corresponding abnormal node frequency, and the specific number of times each node triggers anomalies in bandwidth, latency or number of data packets is counted. , and then make a comprehensive judgment on the entire path based on this. The proportion of each node in the path is taken into account when calculating the trust probability. For example, a probability evaluation algorithm can be constructed according to the frequency of bandwidth anomalies, the proportion of high-value delay segments, and the difference in data packet density. The influence coefficients of node trustworthiness and anomaly frequency are summarized one by one for different paths to obtain the overall trust probability of the path. Next, check whether these trust probabilities are within a reasonable range between 0 and 1. If extreme values ​​that are too high or too low appear, it is necessary to trace back the specific node data to confirm whether there are statistical deviations or omissions, and finally form the trust probability analysis results of each path.

[0160] Based on the statistical analysis results of the credible probability of each attack path, the path credible probability is first distinguished in the range of 0 to 1, and the paths in the range close to 0 are included in the suspected highest risk segment, and the paths in the probability segment greater than 0.5 or higher are classified as medium credible segments. If the credible probability of a path exceeds the high credible threshold determined based on the current network environment and monitoring cycle, it is regarded as a path with relatively acceptable credibility. Then all paths are sorted from low to high. During the comparison process, the number of nodes in the path and the transmission characteristics between nodes are checked at the same time. If more highly abnormal node connections or higher bandwidth occupancy records are found under the same probability, the weight of this path may be lowered during sorting. If most of the nodes in the path are lower than the abnormal threshold established in the previous article, they are given a relatively high priority. Finally, the sorted path list is collected, and the node number, credible probability, number of abnormal records and other information of each path are summarized to generate an attack path priority sequence. .

[0161] The present invention provides a distributed network attack tracing system, comprising:

[0162] The traffic feature analysis module collects the traffic transmission delay, bandwidth occupancy rate and number of data packets between network nodes, and generates node traffic feature clustering values ​​by clustering the data;

[0163] The network relationship building module draws the traffic transmission relationship diagram between nodes based on the node traffic feature clustering value, calculates the link weight and correlation degree, and generates the network topology correlation matrix;

[0164] The traffic sequence analysis module uses the network topology association matrix to track the packet forwarding path and traffic transmission sequence, record the packet identifier and timestamp, and generate the traffic transmission timing sequence;

[0165] The abnormal traffic detection module matches and filters the packet identifiers based on the traffic transmission timing sequence, performs statistical analysis on the traffic transmission characteristics, and generates abnormal traffic transmission indicators;

[0166] The attack path evaluation module combines the abnormal traffic transmission indicators and the node traffic feature clustering values ​​to establish the node credibility scoring standard, calculate the node credibility evaluation value, extract the attack path feature sequence, and generate the attack path priority sequence.

[0167] The above are only preferred embodiments of the present invention and are not intended to limit the present invention in other forms. Any technician familiar with the profession may use the technical contents disclosed above to change or modify them into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention still falls within the protection scope of the technical solution of the present invention.

Claims

1. A distributed network attack tracing method, characterized in that: The following steps are involved: Statistics are collected on the traffic transmission delay, bandwidth occupancy, and number of packets between nodes in the distributed network. The traffic characteristic values ​​of each node are clustered and divided to generate node traffic characteristic clustering values. Based on the node traffic feature clustering values, a traffic transmission relationship diagram between nodes is constructed, link weights and association degrees between nodes are calculated, and a network topology association matrix is ​​generated; According to the network topology association matrix, the packet forwarding path and traffic transmission sequence between nodes are calculated, the packet identifier and timestamp are extracted, and the traffic transmission timing sequence is generated. Based on the traffic transmission timing sequence, the packet identifier is matched and screened, the abnormal traffic transmission characteristics are counted, and the abnormal traffic transmission index is generated; Combining the abnormal traffic transmission index with the node traffic feature clustering value, constructing a node credibility scoring standard, and calculating a node credibility evaluation value; Based on the node credibility evaluation value, extract the attack path feature sequence and generate the attack path feature index; Based on the attack path characteristic index and the node credibility evaluation value, the attack path credibility probability is calculated, the attack paths are prioritized, and an attack path priority sequence is generated.

2. The distributed network attack tracing method according to claim 1 is characterized in that: The steps for obtaining the node traffic feature clustering value are as follows: Statistics are collected on the traffic transmission delay, bandwidth utilization and number of packets between nodes in the distributed network. The traffic transmission data of each node is normalized and the weighted average value of the delay, bandwidth utilization variance and absolute deviation of the number of packets between nodes are calculated to generate a basic data table of traffic characteristics. According to the traffic feature basic data table, feature merging is performed on the weighted average value of the node's delay, the variance of the bandwidth utilization rate, and the absolute deviation of the number of data packets to form a feature data set; According to the characteristic data set, the node traffic characteristic clustering value is calculated, and the formula is: Among them, C k is the node traffic feature clustering value of the kth cluster, m k is the number of nodes in the kth cluster, t i is the weighted average delay of the ith node, b i is the bandwidth utilization variance of the ith node, d i is the absolute deviation of the number of packets of the ith node, is the mean number of packets in the kth cluster.

3. The distributed network attack tracing method according to claim 1 is characterized in that: The steps of obtaining the network topology association matrix are: Based on the node traffic feature clustering value, extract the traffic feature vector between nodes, compare the traffic features between nodes in pairs, calculate the transmission frequency and transmission delay difference between node pairs, and obtain a preliminary node relationship feature data table; According to the preliminary node-to-node relationship feature data table, the link weights between nodes are calculated using the following formula: Among them, W ij is the link weight between node i and node j, F ij is the transmission frequency between node i and node j, T ij is the transmission delay between node i and node j, D i and D j are the traffic feature dimension values ​​of node i and node j respectively, C i and C j are the node traffic feature clustering values ​​of node i and node j respectively; Based on the link weight, the nodes are regarded as points in the graph, the link weight is regarded as the weight of the edge, a traffic transmission relationship graph is constructed, and the graph data structure is converted into a matrix to generate a network topology association matrix.

4. The distributed network attack tracing method according to claim 1 is characterized in that: The steps for obtaining the traffic transmission timing sequence are: According to the network topology association matrix, the link weights in the matrix are traversed one by one, and the forwarding path and forwarding priority of each link are calculated in combination with the size and directionality information of the link weights to form a forwarding path set between nodes; Based on the forwarding path set between the nodes, the path order of the data packet is recorded, the transmission order of the data packet on each path is extracted in turn, and the traffic transmission sequence is generated in combination with the time characteristics of the node link weight; Based on the traffic transmission sequence, a data packet identifier and a timestamp are extracted from the data packet information in the transmission path, and the timing information of all node data packets are integrated in chronological order to generate a traffic transmission timing sequence.

5. The distributed network attack tracing method according to claim 1 is characterized in that: The steps for obtaining the abnormal traffic transmission indicator are as follows: According to the traffic transmission timing sequence, extract the data packet identifiers and the corresponding timestamps one by one, perform statistics on the distribution of the identifiers, record the occurrence frequency and time interval of each identifier, and obtain an identifier time distribution statistics table; Based on the identifier time distribution statistics table, filter the data packet identifiers with abnormal frequency or time interval, combine the timestamps in the traffic transmission timing sequence, extract the time series of abnormal transmission, and form abnormal traffic feature data; Based on the abnormal traffic characteristic data, the abnormal traffic transmission index is calculated, and the calculation formula is: Among them, A is the abnormal traffic transmission index, Q i is the occurrence frequency of the ith packet identifier, CT i is the time interval of the ith packet identifier, E i is the transmission timestamp of the ith packet identifier, is the average value of the transmission timestamps of all abnormal data packets, and k is the total number of abnormal data packet identifiers.

6. The distributed network attack tracing method according to claim 1, characterized in that: The steps for obtaining the node credibility evaluation value are as follows: Combining the abnormal traffic transmission index with the node traffic feature clustering value, statistically analyzing the traffic abnormality performance and traffic feature center value of each node, summarizing the feature deviation of each node, and obtaining a preliminary node abnormal feature data table; Based on the preliminary node abnormal feature data table, analyze the correlation between the node abnormal feature and other nodes, select the node set with greater abnormal feature influence, and obtain node comprehensive abnormal score data; According to the node comprehensive abnormality score data, the node credibility evaluation value is calculated, and the calculation formula is: Among them, R i is the credibility evaluation value of the i-th node, A i is the abnormal traffic transmission index of the i-th node, C i is the node traffic feature clustering value of the i-th node, M i and N i are the abnormal feature dimension value and abnormal correlation degree of the i-th node respectively.

7. The distributed network attack tracing method according to claim 1, characterized in that: The steps for obtaining the attack path characteristic index are as follows: Based on the node credibility evaluation value, nodes whose node credibility evaluation value is lower than the threshold are screened, and the traffic transmission links between the nodes below the threshold are analyzed one by one, and the directionality and data transmission status of each link are extracted, and the transmission correlation degree between the nodes is calculated in combination with the link weight, and a set of traffic transmission links between suspicious nodes is generated; According to the traffic transmission link set between suspicious nodes, the data transmission paths in the links are extracted one by one, the node sequence and data packet transmission rules on the path are analyzed, and the correlation features between the nodes are sorted and classified to extract the attack path feature sequence; Based on the attack path feature sequence, the credibility evaluation value of the node and the transmission link directionality data are called, and the characteristic parameters on each path are counted item by item to generate a feature description of each path to form an attack path feature index.

8. The distributed network attack tracing method according to claim 1, characterized in that: The steps for obtaining the attack path priority sequence are: Based on the attack path characteristic index and the node credibility evaluation value, the node information and characteristic parameters on each attack path are extracted, and weighted combination is performed according to the credibility evaluation value of the nodes in the path, and combined with the abnormal traffic information in the path characteristics, to generate an attack path credibility basic data table; According to the attack path credibility basic data table, the credibility probability of each attack path is calculated, and the influence of the node characteristics and abnormal traffic characteristics of the path on the credibility probability is statistically analyzed item by item to obtain the attack path credibility probability analysis result; Based on the attack path credibility probability analysis result, the attack paths are prioritized from low to high according to credibility probability, the sequence information of the sorted paths is extracted, and the information is integrated to form an attack path priority sequence.

9. A distributed network attack tracing system according to a distributed network attack tracing method according to any one of claims 1 to 8, characterized in that: include: The traffic feature analysis module collects the traffic transmission delay, bandwidth occupancy rate and number of data packets between network nodes, and generates node traffic feature clustering values ​​by clustering the data; The network relationship building module draws the traffic transmission relationship diagram between nodes based on the node traffic feature clustering value, calculates the link weight and correlation degree, and generates the network topology correlation matrix; The traffic sequence analysis module uses the network topology association matrix to track the packet forwarding path and traffic transmission sequence, record the packet identifier and timestamp, and generate the traffic transmission timing sequence; The abnormal traffic detection module matches and filters the packet identifiers based on the traffic transmission timing sequence, performs statistical analysis on the traffic transmission characteristics, and generates abnormal traffic transmission indicators; The attack path evaluation module combines the abnormal traffic transmission indicators and the node traffic feature clustering values ​​to establish the node credibility scoring standard, calculate the node credibility evaluation value, extract the attack path feature sequence, and generate the attack path priority sequence.

Citation Information

Cited By

  • Data stream analysis method based on TCP timestamp detection

    CN120342784A

  • A data flow analysis method based on TCP timestamp detection

    CN120342784B

  • Detection method fusing depth feature extraction and attack recognition

    CN120455178A

  • Network attack detection method and system based on distributed intelligent probe

    CN120896785A