An Abnormal Network Community Detection Method for Dynamic Graphs

Through the abnormal network community detection method for dynamic graphs, the problems of high computing complexity and low detection accuracy in the prior art are solved, and efficient abnormal network community detection is realized to adapt to network changes.

CN119728296BActive Publication Date: 2025-05-30DATA SPACE RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510208738.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-30
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The existing network anomaly detection methods have high computational complexity in large-scale graph structures, low detection accuracy, and it is difficult to effectively detect local anomaly events.

Method used

A method of abnormal network community detection for dynamic graphs is proposed. By obtaining network communication flow data for each time period, dividing nodes into a community set, matching communities in adjacent time periods, calculating the distance of associated community pairs, and determining whether there is an abnormality based on adaptive thresholds.

Benefits of technology

It reduces the computational complexity, improves the accuracy of abnormal network community detection, can effectively detect local abnormal events, and adapt to network changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119728296B_ABST
    Figure CN119728296B_ABST
Patent Text Reader

Abstract

The present invention discloses an abnormal network community detection method for dynamic graphs, which relates to the technical field of anomaly detection and includes: obtaining network communication flow data for each time period, drawing a network communication flow graph, and constructing a network dynamic graph model; dividing the nodes in the network communication flow graph for each time period to obtain a community set; performing community matching based on the network communication flow graphs of adjacent time periods to obtain associated community pairs between adjacent time periods; calculating the distance between the associated community pairs based on the associated community pairs, and determining whether the distance between the associated community pairs is greater than an adaptive threshold, so as to determine whether the community has an anomaly. Compared with monitoring the behavior of a single node or edge in a large-scale network graph structure, the computational complexity is reduced and the detection accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of anomaly detection, and in particular to an anomaly network community detection method for dynamic graphs. Background Art

[0002] Complex networks carry a huge number of nodes and the interaction relationships between nodes. The influx of malicious users makes complex networks face attack threats, bringing huge hidden dangers to network security. Therefore, it is particularly important to detect abnormal behaviors in complex networks in a timely and accurate manner.

[0003] Most of the existing network anomaly detection methods mainly identify abnormal nodes or edges. However, complex networks carry a large number of nodes and complex interaction relationships. Monitoring the behaviors of individual nodes or edges in such a large-scale graph structure has a high computational complexity; at the same time, the current feature extraction methods for network structures are relatively complex, and there is a lack of an efficient similarity measurement method in large-scale datasets. Therefore, when the discrimination between frequent anomalies is small, the accuracy is low; the feature extraction overly relies on global structure information and ignores local details in the network, making it difficult to effectively detect local abnormal events. Summary of the Invention

[0004] In order to overcome the defects of high computational complexity and low detection accuracy in the above-mentioned prior art, the present invention proposes an anomaly network community detection method for dynamic graphs.

[0005] To achieve the above object, the present invention adopts the following technical solutions. An anomaly network community detection method for dynamic graphs includes:

[0006] S1: Obtain the network communication flow data for each time period, draw the network communication flow graph for each time period, and construct a network dynamic graph model, where the network dynamic graph model is composed of network communication flow graphs arranged in chronological order;

[0007] S2. Divide the nodes in the network communication flow graph for each time period to obtain a community set;

[0008] S3. Perform community matching based on the network communication flow graphs of adjacent time periods to obtain associated community pairs between adjacent time periods;

[0009] S4. Calculate the distance between the associated community pairs;

[0010] S5: Determine whether the distance between the associated community pairs is greater than the adaptive threshold; if yes, it means that the community in the latter time period in the associated community pair has an anomaly; if no, it means that the communities in the associated community pair are normal.

[0011] Preferably, in step S2, dividing the nodes in the network communication flow graph for each time period to obtain a community set includes:

[0012] S21: Segment each network communication flow graph to obtain a number of segmented communities;

[0013] S22: Split and merge the segmented communities multiple times to explore all possible ways of partitioning the communities;

[0014] S23: Calculate the modularity density for each partitioning method, find the partitioning method with the maximum modularity density, and mark the several non - overlapping communities formed by the partitioning method with the maximum modularity density as the optimal non - overlapping communities, thereby forming a community set , where, C t is the community set formed in the t - th time period, represents the q - th community in the t - th time period, and m is the total number of communities.

[0015] Preferably, the calculation process of the modularity density includes:

[0016] C1: For each partitioning method, obtain the modularity Q based on multiple communities. The calculation formula for the modularity Q is:

[0017] ;

[0018] where, C is the community set, c i is the i - th community in C, is the number of edges between nodes within the i - th community, is the number of edges from nodes within the i - th community to nodes outside the i - th community, and |E| is the total number of edges in the network communication flow graph;

[0019] C2: Calculate the splitting penalty SP based on multiple communities. The calculation formula for the splitting penalty SP is:

[0020] ;

[0021] where, represents the number of edges connecting the i - th community and the j - th community;

[0022] C3: Obtain the modularity density Q ds based on the modularity and the splitting penalty. The calculation formula for the modularity density Q ds is:

[0023] ;

[0024] where, is the density within the i - th community; is the density between the i - th community and the j - th community.

[0025] Preferably, in step S3, the community matching based on the network communication flow graphs of adjacent time periods to obtain the associated community pairs between adjacent time periods includes:

[0026] S31: Based on two communities of adjacent time periods , , calculate the similarity J pq between the two communities, and the formula is:

[0027] ;

[0028] Among them, is the number of overlapping nodes of community and community , is the sum of the number of nodes of community and community ; is the p-th community in the (t - 1)-th time period, is the q-th community in the t-th time period;

[0029] S32: Compare the similarity between the two communities with a set threshold, obtain several communities in the t-th time period corresponding to the p-th community in the (t - 1)-th time period that exceed the set threshold, and mark them as communities to be selected;

[0030] S33: Find the community with the highest similarity between the communities to be selected and the p-th community in the (t - 1)-th time period as the community associated with the p-th community in the (t - 1)-th time period, so as to obtain the associated community pairs between adjacent time periods .

[0031] Preferably, in step S4, calculating the distance of the associated community pairs includes:

[0032] S41: Based on all the matched associated community pairs , obtain the Laplacian matrix and ;

[0033] S42: Perform feature extraction on the Laplacian matrix and to obtain the NetLSD vectors and ;

[0034] S43: Based on the NetLSD vectors and to obtain the distance d t of the associated community pairs;

[0035] Among them, is the p-th community in the (t-1)-th time period, is the q-th community in the t-th time period; and are the Laplacian matrices in the (t-1)-th time period and the t-th time period respectively; and are respectively and network Laplacian spectrum description vectors; d t is the distance of the associated community pair .

[0036] Preferably, in step S5, the obtaining step of the adaptive threshold includes:

[0037] SS1: Based on the distance d t-1 of the associated community pair between the (t-2)-th time period and the (t-1)-th time period, obtain the distance prediction value of the associated community pair between the (t-1)-th time period and the t-th time period, and the calculation formula is:

[0038] ;

[0039] where α is a smoothing coefficient and 0 < α < 1, is the distance prediction value of the associated community pair between the (t-2)-th time period and the (t-1)-th time period;

[0040] SS2: Calculate the distance mean avg and the distance standard deviation V of the associated community pairs within the sliding window, and the calculation formulas are respectively:

[0041] ;

[0042] where k is the size of the sliding window, t is the time period number, and d x is the distance of the associated community pair between the (x-1)-th time period and the x-th time period, x = t - k + 1,..., t;

[0043] SS3: Based on the distance prediction value of the associated community pair and the distance standard deviation V of the associated community pairs within the sliding window, obtain the adaptive threshold U t , and the calculation formula is: ;

[0044] where β is the standard deviation coefficient.

[0045] Preferably, the calculation formulas of the and are respectively:

[0046] ;

[0047] where |c i | is the number of nodes in the i-th community.

[0048] Preferably, in step S1, segmentation is performed at fixed time intervals to form multiple time periods. In each time period, the IP of each host in the network communication is abstracted as a node, and the network communication between nodes is abstracted as an edge, thereby forming a network communication flow graph.

[0049] Preferably, in step S21, the network communication flow graph for each time period is segmented to obtain a number of segmented communities, and each community contains at least one node.

[0050] Preferably, the distance d t between associated community pairs is calculated by the formula:

[0051] ;

[0052] where represents the 2-norm operation.

[0053] The advantages of the present invention are as follows:

[0054] (1) By partitioning the nodes in the network communication flow graph, the present invention obtains a number of optimal non-overlapping communities. Then, based on the communities in different time periods, associated community pairs between adjacent time periods are obtained. Based on the associated community pairs, it is identified whether there are abnormal network communities. Compared with monitoring the behavior of individual nodes or edges in a large-scale network graph structure, the computational complexity is reduced, and the accuracy of abnormal network community detection is improved.

[0055] (2) By segmenting the network communication flow graph for each time period, the present invention obtains a number of segmented communities; performs multiple splits and merges on the segmented communities, explores all possible partitioning methods, and calculates the modularity density for each partitioning method. The number of non-overlapping communities formed by the partitioning method with the maximum modularity density is marked as the optimal non-overlapping communities, which can prevent the community partition from being too small and avoid the resolution limit problem, ensuring the scientificity and rationality of the community partition and making the subsequent abnormal detection of the network community more accurate.

[0056] (3) Through the similar overlapping nodes between communities in different time periods, the present invention calculates the similarity between two communities, obtains the communities to be selected based on the similarity, and searches for the community with the highest similarity between the community to be selected and the p-th community in the (t - 1)-th time period as the q-th community in the t-th time period associated with the p-th community in the (t - 1)-th time period, thereby obtaining the associated community pairs , by calculating the similarity of all possible community pairs in adjacent time periods, it is possible to identify which community in the next time period is most similar to the community in the previous time period, thereby tracking the evolutionary path of the community over time and having high search efficiency.

[0057] (4) Based on the distance between associated community pairs, the present invention obtains the average distance avg and the distance standard deviation V of the associated community pairs within the sliding window. Based on the predicted value of the associated community distance and the distance standard deviation of the associated community pairs within the sliding window, the adaptive threshold is dynamically adjusted. Compared with the fixed threshold, which may ignore the dynamics of network traffic and cause false alarms or missed alarms, the adaptive threshold is dynamically adjusted according to the normal fluctuation range of the current network traffic, ensuring the rationality and real-time nature of the threshold setting, and effectively improving the accuracy and robustness of anomaly detection.

[0058] (5) By calculating the distance between associated community pairs and combining with the adaptive threshold, the present invention can effectively detect abnormal communities in the network. The introduction of the adaptive threshold enables the method to adapt to changes in the network and improve the detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is a network anomaly comparison diagram;

[0060] Figure 2 is a process diagram of the formation of a network communication flow graph;

[0061] Figure 3 is a construction diagram of community discovery for adjacent time periods;

[0062] Figure 4 is a structural flowchart of the method of the present invention;

[0063] Figure 5 is a flowchart of the steps of the method of the present invention;

[0064] Figure 6 is a comparison diagram of the results of the method of the present invention and the NetLSD method. DETAILED DESCRIPTION OF THE INVENTION

[0065] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0066] As Figures 1-5 shown, the present invention proposes an abnormal network community detection method for dynamic graphs, including:

[0067] S1: Obtain network communication flow data, draw a network communication flow graph, and construct a network dynamic graph model, where the network dynamic graph model is composed of network communication flow graphs arranged in chronological order.

[0068] In network communication, a host generates associations with other hosts by sending and receiving data packets. The network communication flow data exchanged between hosts is usually defined as a series of data packets with the same network triple, including information such as the source host, destination host, and packet size.

[0069] Divide it at fixed time intervals to form multiple time periods. In each time period, abstract the IP of each host in the network communication into a node, and abstract the network communication between node v i and node v j into an edge e ij , thus forming a network communication flow graph , where v and ε represent the sets of nodes and edges respectively.

[0070] Based on the network communication flow graphs of multiple time periods of the network communication flow, construct a network dynamic graph model , represents the network communication flow graph in the t-th time period, T is the last time period, and e=(i,j,w)∈ε t means that at the t-th time period, the i-th node and the j-th node are connected in network communication and the weight is w. w represents the tightness of the network communication connection between node i and node j. For an unweighted graph, w is always 1, and for a weighted graph, w∈R + .

[0071] Because when extracting the graph structure of the network from the network communication flow data, since it is impossible to use the entire data set at once to create a unique static overall graph, the data set needs to be divided at fixed time intervals. As Figure 2 shown, times such as t 1 , t 2 , t 3 etc. are divided into one time period. Using the network communication flow data at these times, a network communication flow graph can be constructed.

[0072] Next, construct a dynamic graph model through the network communication flow graphs of multiple time periods, and represent the dynamic graph model as , which is equivalent to splicing multiple static network sequences to form a dynamic network.

[0073] In this embodiment, for step S1, the fixed time interval is set to 5 minutes.

[0074] S2: Divide the nodes in the network communication flow graph of each time period to obtain a community set.

[0075] A subtle anomaly refers to a deviation or irregularity that occurs within a small area of densely connected nodes in network communication, i.e., within a community. These deviations or irregularities are not easily detectable by global analysis methods. Figure 1 In a graph of the global structure, it is difficult to observe that three malicious users are committing fraud. Therefore, the present invention proposes an abnormal network community detection method for dynamic graphs.

[0076] To enhance the detection ability of network subtle anomalies, features are extracted from groups (i.e., communities) rather than from individual nodes, edges, or the entire network communication flow graph. Here, the time dimension and network structure need to be considered. First, all community information needs to be extracted based on the given time series data, which requires the use of community discovery algorithms. As time changes, the composition structure of the community changes, and the community information for each time period needs to be corresponded. At this time, community matching technology is required. After determining the matching communities, an anomaly detection method is used to compare the changes in associated communities in two adjacent time periods to determine whether there is a possibility of anomalies in the community.

[0077] In this embodiment, the Louvain community discovery algorithm is adopted to obtain a community set, including:

[0078] S21: Segment the network communication flow graph for each time period to obtain a number of divided communities;

[0079] S22: Split and merge the divided communities multiple times to explore all possible ways of dividing the communities;

[0080] S23: Calculate the modularity density for each division method, find the division method when the modularity density is the largest, and mark the number of non-overlapping communities formed by the division method with the largest modularity density as the optimal non-overlapping communities, thereby forming a community set , where C t is the community set formed by the t-th time period, which is obtained by splitting the node set v in the t-th time period t , represents the q-th community in the t-th time period, and m is the total number of communities.

[0081] The calculation method of the modularity density includes:

[0082] C1: For each division method, obtain the modularity Q based on multiple communities. The calculation formula of the modularity Q is:

[0083] ;

[0084] Among them, C is the set of communities, and c i is the i-th community in C, is the number of edges between nodes within the i-th community, is the number of edges from the nodes within the i-th community to the nodes outside the i-th community, and |E| is the total number of edges in the network communication flow graph.

[0085] The modularity Q is an indicator used to evaluate the quality of community detection, comparing the actual number of edges within a community with the expected number of edges in a random case with the same number of nodes and degree sequence. The difference between these two values forms the basis of the modularity measure. The overall idea is that the divided communities should exhibit more and tighter connections than in the random case. Therefore, a Q value close to 0 means that the result of this community detection is close to random assignment. The higher the Q value, the stronger the community structure.

[0086] C2: Calculate the splitting penalty SP based on multiple communities. The calculation formula for the splitting penalty SP is:

[0087] ;

[0088] Among them, represents the number of edges connecting the i-th community and the j-th community.

[0089] C3: Obtain the modularity density Q based on modularity and splitting penalty ds , and the modularity density Q ds has the following calculation formula: ;

[0090] Among them, is the density within the i-th community; is the density between the i-th community and the j-th community, which is the number of edges connecting the i-th community and the j-th community divided by the product of the number of nodes in the two communities.

[0091] and have the following calculation formulas respectively:

[0092] ;

[0093] Among them, |c i | is the number of nodes in the i-th community.

[0094] However, modularity has some limitations. It tends to merge small communities into large communities and cannot effectively identify smaller communities, which is called the resolution limit problem at this time. To alleviate this drawback, the splitting penalty and community density are mixed into modularity to obtain the modularity density, and the modularity density is used to find the optimal community partitioning method. At the same time, the splitting penalty can prevent the problem that the communities obtained by community partitioning are too small.

[0095] The divided communities can be a set of communities each containing one node, where each community in this set of communities covers different nodes in the network, or several pre - divided communities obtained by other technologies, and then optimized by the method of the present invention to obtain several non - overlapping communities formed by the partitioning method with the maximum modularity density.

[0096] The goal of community discovery is to identify groups of closely - connected nodes in the network and regard the group of closely - connected nodes as the same community. As Figure 3 shown, in each time period, the node set v of the network communication flow graph t is segmented into several non - overlapping communities, so as to deeply understand the group behavior and dynamics in the network.

[0097] S3: Perform community matching based on the network communication flow graphs of adjacent time periods to obtain pairs of associated communities between adjacent time periods.

[0098] After community discovery, the nodes in the network communication flow graph are assigned to communities. However, the numbers or labels of communities in different time periods may not correspond. Community matching technology associates the corresponding communities in adjacent time periods to understand their similarities and differences.

[0099] Let and represent the sets of communities formed by the network communication flow graph in the (t - 1) - th time period and the network communication flow graph in the t - th time period respectively, where and represent the p - th community of and the q - th community of . Community matching technology obtains pairs of associated communities

[0100] based on the similar overlapping nodes between communities in different time periods.

[0101] S31: Based on two communities , in adjacent time periods, calculate the similarity J pq between the two communities. The formula is:

[0102] ;

[0103] where is the community and the community The number of overlapping nodes, is the sum of the number of nodes in the community and the community; is the p-th community in the (t - 1)-th time period, and is the q-th community in the t-th time period;

[0104] S32: Compare the similarity between two communities with a set threshold, obtain several communities in the t-th time period corresponding to the p-th community in the (t - 1)-th time period that exceed the set threshold, and mark them as communities to be selected;

[0105] S33: Find the community with the highest similarity between the communities to be selected and the p-th community in the (t - 1)-th time period as the community in the t-th time period associated with the p-th community in the (t - 1)-th time period, so as to obtain the associated community pairs between adjacent time periods .

[0106] For any community in the (t - 1)-th time period , check its similarity with all communities in the t-th time period . If this similarity value exceeds the set threshold τ, it is considered that the two communities have a corresponding relationship in the time series. When the similarities of multiple communities all exceed this threshold, select the community with the highest similarity as the matching community for adjacent time periods. The similarity value ranges from 0 to 1, where 0 means that the two communities in adjacent time periods have no common nodes and no corresponding relationship, and 1 means that the nodes in the two communities are the same. By calculating the similarities of all possible community pairs between adjacent time periods, it is possible to identify which community in the t-th time period is most similar to the community in the (t - 1)-th time period, thereby tracking the evolution path of the community over time.

[0107] S4: Calculate the distance between the associated community pairs, including:

[0108] S41: Based on all matching associated community pairs , obtain the Laplacian matrix and ;

[0109] S42: Perform feature extraction on the Laplacian matrix and to obtain the NetLSD vectors and ;

[0110] S43: Based on the NetLSD vectors and obtain the distance d of the associated community pairs t ;

[0111] The distance d between associated community pairs t The calculation formula is as follows:

[0112] ;

[0113] Wherein, and are the Laplacian matrices of the (t - 1)-th time period and the t-th time period respectively, and are respectively and Network Laplacian Spectral Descriptor (NetLSD), which is composed of the traces of heat kernel matrices at different times, and the heat kernel matrix is obtained by exponentiating the Laplacian eigenvalues; d t is the distance between the associated community pair ; represents the 2-norm operation. This graph processing method can better reflect the global information of the graph over time.

[0114] S5: Determine whether the distance d between the associated community pairs t is greater than the adaptive threshold U t . If so, mark the community in the t-th time period, i.e., has an anomaly; if not, it means that the community in the t-th time period is normal.

[0115] Under normal network conditions, the differences between corresponding communities in adjacent time periods are very small. This means that the lower the value of the distance d t between the associated community pairs, the closer the network state is to normal. When the network is under abnormal attack, the similarity between the abnormal and normal network graphs will decrease significantly, resulting in a sharp increase in the structural distance, and the increase in the distance d t between the associated community pairs indicates an increased possibility of an attack behavior in the network. The network state changes dynamically over time. The normal network state is not static but will be updated with changes in various factors such as the network environment. Therefore, network communication has dynamics and volatility. Using a fixed threshold to detect network abnormal states may lead to a decrease in detection accuracy. The dynamic threshold will be adjusted dynamically according to the immediate changes and volatility characteristics of the network state, which not only avoids the accuracy problems that may be caused by static thresholds but also can be effectively adjusted based on the range of normal state fluctuations. Therefore, it is crucial to define the network threshold in combination with the dynamic changes of the network structure. We detect whether the network is under attack by applying the sliding window technique and the adaptive threshold method.

[0116] The steps for obtaining the adaptive threshold U t include:

[0117] SS1: Obtain the predicted value of the distance between the associated community pairs in the \(t\)-th time period based on the distance \(d\) between the associated community pairs in the \((t - 2)\)-th time period and the \((t - 1)\)-th time period t-1 , and the calculation formula is: , and the calculation formula is:

[0118] ;

[0119] where \(\alpha\) is a smoothing coefficient and \(0\lt\alpha\lt1\), is the predicted value of the distance between the associated community pairs in the \((t - 2)\)-th time period and the \((t - 1)\)-th time period.

[0120] SS2: Calculate the average value \(avg\) and the standard deviation \(V\) of the distances between the associated community pairs within the sliding window based on the distances between the associated community pairs. The calculation formulas are respectively:

[0121] ;

[0122] where \(k\) is the size of the sliding window, \(t\) is the time period number, and \(d\) x is the distance between the associated community pairs in the \((x - 1)\)-th time period and the \(x\)-th time period, and \(x=t - k + 1,\cdots,t\).

[0123] SS3: Obtain the adaptive threshold \(U\) based on the predicted value of the distance between the associated community pairs and the standard deviation \(V\) of the distances between the associated community pairs within the sliding window t , and the calculation formula is: .

[0124] where \(\beta\) is the standard deviation coefficient.

[0125] Compared with the fixed threshold, which may ignore the dynamics of network traffic and cause false alarms or missed alarms, the adaptive threshold is dynamically adjusted according to the normal fluctuation range of the current network traffic, ensuring the rationality and real-time nature of the threshold setting, and effectively improving the accuracy and reliability of anomaly detection.

[0126] At the same time, the present invention adopts the exponentially weighted moving average model and the sliding window technology, which can accurately capture the dynamic changes of the network state and improve the detection accuracy. By precisely setting the dynamic threshold, the false positive rate is further reduced.

[0127] In this embodiment, the network status data of the Vast challenge 2013 is selected as the experimental data set, and the data volume is about 5GB. The hardware experimental environment is as follows: the processor is an Intel(R) Core(TM) i7-12700K CPU @ 3.6GHz, and the graphics card is an NVIDIA GeForce GTX 1660 GPU. The software environment is Python 3.8.

[0128] In this embodiment, the sliding window is set to k = 5, and the standard deviation coefficient is set to β = 1.5. First, using the community discovery algorithm in the present invention, community partitioning and community matching are performed on the above data set. After finding community No. 1, the anomaly detection accuracy of community 1 is tested under different values of the sliding window k and the standard deviation coefficient β. The range of the sliding window size can be set to [5, 20], increasing by 5 each time; the range of the standard deviation coefficient size is set to [0.5, 2], increasing by 0.5 each time. According to different sizes of the sliding window and the standard deviation coefficient, and according to step S5 and the definition of anomaly detection accuracy, the anomaly detection accuracy results for community 1 can be obtained as shown in the following table:

[0129] Detection results of community 1 under different parameter settings

[0130] ;

[0131] It can be seen that for community 1, when the sliding window is set to k = 5 and the standard deviation coefficient is set to β = 1.5, the detection effect of community 1 is the best. Therefore, the sliding window is set to k = 5, and the standard deviation coefficient is set to β = 1.5. Then, the four detection indicators (F1 score, recall rate, precision rate, accuracy rate) of all communities are calculated and compared with the NetLSD method, as Figure 6 shown.

[0132] From Figure 6 it can be seen that the method in the present invention performs anomaly detection on the network status data set of the Vast challenge 2013. The community with the best detection effect and the average value of the detection results of all communities are better than the detection results obtained using the NetLSD method, and the effect comparison is significant in terms of accuracy.

[0133] Of course, for those skilled in the art, the present invention is not limited to the details of the above-described exemplary embodiments, but also includes the same or similar structures that can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims concerned.

[0134] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

[0135] The technologies, shapes, and structures not described in detail in the present invention are all well-known technologies.

Claims

1. A method for detecting abnormal network communities in dynamic graphs, characterized in that: include: S1: Obtain network communication flow data for each time period, draw a network communication flow graph for each time period, and construct a network dynamic graph model, wherein the network dynamic graph model is composed of network communication flow graphs arranged in order of time periods; S2, divide the nodes in the network communication flow graph of each time period to obtain the community set, including: S21: Segment each network communication flow graph to obtain a number of segmented communities; S22: split and merge the divided communities multiple times to explore all possible ways of dividing the communities; S23: Calculate the modularity density for each partitioning method, find the partitioning method with the maximum modularity density, and mark the non-overlapping communities formed by the partitioning method with the maximum modularity density as the optimal non-overlapping communities, thereby forming a community set. , where C t is the community set formed in the tth time period, represents the qth community in the tth time period, and m is the total number of communities; S3, performing community matching based on the network communication flow graph of adjacent time periods, and obtaining the associated community pairs between adjacent time periods, including: S31: Two communities based on adjacent time periods , , calculate the similarity J between two communities pq , the formula is: ; in, For the community and community The number of overlapping nodes, For the community With the community The sum of the number of nodes; is the pth community in the t-1th time period, is the qth community in the tth time period; S32: Compare the similarity between the two communities with a set threshold, obtain several communities in the t-th time period corresponding to the p-th community in the t-1-th time period that exceed the set threshold, and mark them as communities to be selected; S33: Find the community with the highest similarity with the pth community in the t-1th time period among the communities to be selected, and use it as the community associated with the pth community in the t-1th time period, so as to obtain the associated community pairs between adjacent time periods. ; S4, calculate the distances of associated community pairs; S5: Determine whether the distance between the associated community pairs is greater than the adaptive threshold; if so, it indicates that an abnormality occurs in the community in the latter time period of the associated community pair; if not, it indicates that the community in the associated community pair is normal.

2. The abnormal network community detection method for dynamic graphs according to claim 1, characterized in that: The calculation process of the modularity density includes: C1: For each division method, the modularity Q is obtained based on multiple communities. The calculation formula of modularity Q is: ; Among them, C is the community set, c i is the i-th community in C, is the number of edges between nodes in the ith community, is the number of edges from nodes in the ith community to nodes outside the ith community, |E| is the total number of edges in the network communication flow graph; C2: Calculate the segmentation penalty SP based on multiple communities. The calculation formula of the segmentation penalty SP is: ; in, represents the number of edges connecting the i-th community and the j-th community; C3: Obtain modularity density Q based on modularity and segmentation penalty ds , modularity density Q ds The calculation formula is: ; in, is the density within the i-th community; is the density between the i-th community and the j-th community.

3. The abnormal network community detection method for dynamic graphs according to claim 1, characterized in that: In step S4, the distance between the associated community pairs is calculated, including: S41: Associated communities based on all matches , get the Laplace matrix and ; S42: Laplacian Matrix and Perform feature extraction to obtain the NetLSD vector and ; S43: Based on NetLSD vector and Get the distance d between the associated community pairs t ; in, is the pth community in the t-1th time period, is the qth community in the tth time period; and are the Laplace matrices of the t-1th time period and the tth time period respectively; and They are and The network Laplace spectrum describes the vector; d t For the associated community ) distance.

4. The abnormal network community detection method for dynamic graphs according to claim 1, characterized in that: In step S5, the step of obtaining the adaptive threshold comprises: SS1: Based on the distance d between the associated community pairs in the t-2th time period and the t-1th time period t-1 , get the distance prediction value of the associated community pair between the t-1th time period and the tth time period , the calculation formula is: ; Where α is the smoothing coefficient and 0<α<1, is the predicted distance value of the associated community pair between the t-2th time period and the t-1th time period; SS2: Calculate the distance mean avg and distance standard deviation V of the associated community pairs in the sliding window. The calculation formulas are: ; Among them, k is the size of the sliding window, t is the time period number, and d x is the distance between the associated community pairs in the x-1th time period and the xth time period, x=t-k+1,...,t; SS3: Obtain the adaptive threshold U based on the distance prediction value of the associated community pair and the distance standard deviation V of the associated community pair in the sliding window t , the calculation formula is: ; Among them, β is the standard deviation coefficient.

5. The abnormal network community detection method for dynamic graphs according to claim 2, characterized in that: Said and The calculation formulas are: ; Among them, |c i | is the number of nodes in the i-th community.

6. The abnormal network community detection method for dynamic graphs according to claim 1, characterized in that: In step S1, the network is divided into multiple time periods according to fixed time intervals. In each time period, the IP of each host in the network communication is abstracted as a node, and the network communication between nodes is abstracted as an edge, thereby forming a network communication flow graph.

7. The abnormal network community detection method for dynamic graphs according to claim 1, characterized in that: In step S21, the network communication flow graph of each time period is segmented to obtain a plurality of segmented communities, each of which contains at least one node.

8. The abnormal network community detection method for dynamic graphs according to claim 3, characterized in that: The distance d between the associated community pairs t The calculation formula is: ; in, Represents the 2-norm operation.

Citation Information

Patent Citations

  • Overlapping community identification method and device, equipment, storage medium and program product

    CN114329099A

  • Dynamic graph anomaly detection method based on community structure

    CN114443909A