Social network abnormal user detection method based on graph flow triangle counting

By adopting distributed architecture and real-time sampling algorithms in social networks, the problems of high memory usage, reduced sampling rate and difficult data division coordination in the prior art are solved, and more efficient and accurate abnormal user detection effects are achieved.

CN119988733APending Publication Date: 2025-05-13NORTHEASTERN UNIV CHINA

Patent Information

Application Number
CN202510080472.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When the prior art efficiently detects abnormal users in social networks, it faces the problems of high memory usage, reduced sampling rate, and difficulty in data division and calculation coordination in a distributed environment, which affects the accuracy and efficiency of triangle counting.

Method used

The social network anomaly user detection method based on graph flow triangle counting is adopted. Through distributed architecture and real-time sampling algorithm, the estimation accuracy of triangle counting and sensitivity to local network structure changes are improved, and the communication protocol is optimized to reduce resource consumption.

Benefits of technology

It realizes more accurate abnormal user detection, and improves the performance and applicability of the algorithm, especially in complex distributed environments with high throughput and low latency requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988733A_ABST
    Figure CN119988733A_ABST
Patent Text Reader

Abstract

The invention provides a method for detecting abnormal users of a social network based on graph flow triangle counting, relates to the technical field of abnormal detection of the social network, and adopts an approximate calculation technology based on sampling to efficiently complete triangle counting of dynamic graph flow under the constraint of a limited memory. The real-time performance and accuracy of triangle counting are remarkably improved, and the real-time processing requirement of large-scale graph flow is met; according to the method, by dynamically calculating the clustering coefficient of the nodes and combining the abnormal behavior characteristics, the abnormal vertexes in the social network can be effectively identified, and the method has higher accuracy in the aspect of detecting abnormal users and abnormal social modes; according to the method, efficient distribution and parallel execution of calculation tasks are achieved through distributed calculation, the bottlenecks of single-machine memory and processing capacity are broken through, the real-time estimation precision of the algorithm is improved, and meanwhile the communication overhead is further reduced through load balancing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of social network anomaly detection, and in particular to a method for detecting abnormal users in social networks based on graph flow triangle counting. Background Art

[0002] With the advent of the big data era, graph-structured data has been widely used in social network analysis because it can effectively express complex relationships and structures. The interactions between users in social networks can be naturally modeled as graphs, where nodes represent users and edges represent relationships or interactions between users. By processing and analyzing graph-structured data, user behavior patterns and structural characteristics of the network can be mined. However, with the rapid growth and dynamic changes in the scale of social networks, how to efficiently detect abnormal users has become an important research topic.

[0003] As a core problem in graph analysis, triangle counting is of great significance in social networks. Real social networks often contain a large number of triangles, which reflects the strong correlation and community structure between users, while abnormal users usually have significantly different interaction patterns from other users, and are often characterized by fewer participating triangles or abnormal distribution. By counting the number of triangles, we can further obtain the clustering coefficient, which measures the local connection density of nodes and is an important indicator for detecting abnormal users. For example, users with abnormally low clustering coefficients may have different behavior patterns from the main body of the network and may be isolated fake accounts or malicious users.

[0004] In the graph flow model, the edges and nodes of social networks are constantly arriving or changing dynamically, making it infeasible to store complete historical graph data. With limited memory space, how to efficiently perform approximate triangle counting and calculate the clustering coefficient has become a hot topic in current research. Sampling technology has been widely used in the study of graph flow triangle counting due to its efficiency and scalability. In recent years, abnormal user detection methods based on graph flow triangle counting have gradually attracted attention. By using the number of triangles and clustering coefficients of nodes, potential fake accounts, malicious behaviors, and abnormal social patterns can be identified, thereby improving the security and credibility of social networks.

[0005] However, this approach also faces some challenges. For example, in order to ensure that the sampling rate is not lower than the threshold, a certain amount of memory resources are required. As the scale of the graph stream continues to expand, under fixed memory constraints, the sampling rate may gradually decrease, affecting the accuracy of triangle counting. In addition, distributed computing provides a possible solution for the processing of large-scale graph streams. By distributing graph data to multiple computing nodes for parallel processing, the efficiency and scalability of the algorithm can be significantly improved. However, in a distributed environment, how to effectively divide the graph data, coordinate the calculation results of each node, and maintain the accuracy of triangle counting is still an important problem that needs to be solved.

[0006] A Chinese patent with application number 202310264542.9 discloses "a method, apparatus and computer device for detecting abnormal behavior in a social network". The method stores persistent subgraph patterns and non-persistent subgraph patterns separately, wherein the counting slot only stores the persistent cumulative value of the non-persistent subgraph without dividing the non-persistent subgraph pattern. The persistent subgraph pattern is divided and the persistent value count is performed in the storage bucket. This reasonable allocation of memory greatly improves detection efficiency and reduces memory and computing costs. At the same time, the persistent subgraph pattern will be updated in real time at each timestamp to ensure that the memory is reasonably allocated.

[0007] The Chinese patent with application number 202110777154.1 discloses a "triangle counting method and device for online adaptive sampling space of large-scale stream graphs". The method is that when the scale of the data graph stream is unknown and there is sufficient available memory, when a new edge arrives, the reservoir space can be incrementally and adaptively maintained through two samplings of the reservoir, ensuring that the real-time sampling rate is not lower than the threshold given by the user, so that more triangles can be found, thereby improving the accuracy of triangle number evaluation.

[0008] The method of the Chinese patent with application number 202310264542.9 requires converting the social graph stream data into a social network snapshot graph, and then performing abnormal behavior detection on the snapshot graph. Since the snapshot graph often contains a large number of nodes and edges, its scale may far exceed the memory limit. Directly calculating the number of triangles and clustering coefficients of all nodes faces memory bottlenecks and computing efficiency issues.

[0009] The method of the Chinese patent with application number 202110777154.1 requires sufficient memory to ensure that the sampling rate is not lower than the threshold. As the scale of graph stream data continues to expand, under the premise of fixed memory usage, the sampling rate will inevitably decrease gradually, thereby affecting the accuracy and evaluation effect of triangle counting results. This performance degradation phenomenon is particularly significant when processing large-scale dynamic graph data. Summary of the invention

[0010] In view of the deficiencies in the prior art, the present invention provides a method for detecting abnormal users in social networks based on graph stream triangle counting, which is used to efficiently detect abnormal users in social networks in the form of graph stream data.

[0011] The technical solution of the present invention is as follows:

[0012] A method for detecting abnormal users in social networks based on graph flow triangle counting includes the following steps:

[0013] Step 1: Obtain a graph stream dataset that reflects the relationships between users in a real-world social network;

[0014] Step 2: Preprocess the graph stream data set and randomly select a portion of the data as the training set;

[0015] The preprocessing is specifically as follows: converting each piece of data in the graph stream data set into a triple form containing two vertices and time information, i.e., edge stream e = (u, v, t), where u and v are vertices, t is a timestamp, vertices represent users in a social network, and edges represent network interactions between users; then, after the current timestamp of the graph stream data set exceeds 50% of the total time, randomly inserting abnormal vertex information to simulate the interaction behavior of abnormal users in the network;

[0016] Step 3: Calculate the triangle count and vertex degree according to the training set, combine them to get the vertex weight value, and then intercept the first n% of the vertices with the largest weight as the weight information table. The weight value calculation formula is:

[0017] W e =λ 1 ×η+λ 2 (δ 1 +δ 2 ) (1)

[0018] where η represents the local triangle count of edge e = (u, v), δ 1 and δ 2 represents the degree of the edge e = (u, v) between two vertices u and v, λ 1 and λ 2 represents the weight coefficient;

[0019] Step 4: Distribute the training set from the master node to the worker nodes through a distributed architecture;

[0020] Step 4.1, set the hash function according to the number of working nodes, and calculate the hash values ​​of the two vertices u and v of the edge flow;

[0021] Step 4.2: Add a tag to the edge flow according to the hash value of u and v and the load of the working node;

[0022] Step 4-3, send the marked edge flow to the working node;

[0023] The distributed architecture specifically adopts the three-layer architecture of Master-Worker-Aggregator; the master node Master is used to allocate tasks and coordinate the computing work of each worker node; the worker node Worker is used to count global and local triangles and process graph stream data locally; the aggregation node Aggregator is used to summarize the counting results of each worker node;

[0024] The distributed architecture adopts unicast and broadcast communication forms, and the hash function distributes the edge flow in combination with the load situation; if the load situation of the working node does not exceed the set threshold, the distribution situation is determined according to the vertex hash value; if the load situation of the working node exceeds the threshold, the vertex is preferentially distributed to the working node with the smallest load;

[0025] Step 5: Real-time estimation of global triangle and local triangle counts through working nodes;

[0026] Step 5.1, initialize the working node, set the sampling set to empty, and set the counter to 0;

[0027] The sampling set is composed of a waiting queue, a weighted queue and a reservoir pool in sequence, and the proportion of the three types of memory occupied is controlled by parameters, and the parameters specifically include the total capacity k of the sampling set, the waiting queue ratio α and the weighted queue ratio β; the waiting queue capacity is W=αk, the weighted queue capacity is H=(1-α)k, and the reservoir pool capacity is S=(1-α)(1-β)k;

[0028] Step 5.2, receiving edge streams and updating real-time estimates of global triangle and local triangle counts;

[0029] Among them, the real-time estimate of the current triangle count needs to be calculated based on the edge flow. The specific situation includes:

[0030] (1) If both edge u and edge v are in the reservoir, the sampling probability of the triangle is:

[0031]

[0032] (2) If one of the edges u or v is in the reservoir pool and the other is in the waiting queue or weight queue, the sampling probability of the triangle is:

[0033]

[0034] (3) If both edge u and edge v are not in the reservoir pool, the sampling probability of the triangle is:

[0035] p uvw =1 (4)

[0036] Among them, l is the total number of edges in the current graph flow.

[0037] Step 5.3: Determine whether to perform sampling according to the edge flow marking situation;

[0038] By judging whether the mark of the edge flow is equal to the sequence number of the working node, if so, the working node needs to perform sampling and execute step 5.4; if not, the edge flow is discarded and step 6 is executed;

[0039] Step 5.4: Sample the edge flow according to the weight information table, update the sampling set, and sample the sampling set according to the preset parameters and capacity. The specific steps are as follows:

[0040] Step 5.4.1: The waiting queue is a first-in-first-out queue. After receiving the edge flow, the available capacity is determined. If there is available capacity, step 5.4.2 is executed; if there is no available capacity, step 5.4.3 is executed.

[0041] Step 5.4.2, queue the edge flow and execute step 6;

[0042] Step 5.4.3, add the edge into the queue, dequeue the edge at the end of the queue, and send the dequeued edge to the weighted queue, and execute step 5.4.4;

[0043] Step 5.4.4: The weighted queue is a priority queue. After receiving the outgoing edge, a weight value is assigned to the outgoing edge according to the weight information prediction table to determine the available capacity. If there is available capacity, execute step 5.4.5; if there is no available capacity, execute step 5.4.6;

[0044] Step 5.4.5, queue the outgoing edge and execute step 6;

[0045] Step 5.4.6, queue the outgoing edge, dequeue the edge at the end of the queue, and send the outgoing edge to the reservoir pool, and execute step 5.4.7;

[0046] Step 5.4.7, the reservoir pool adopts equal probability sampling. After receiving the outbound edge, it determines the available capacity. If there is available capacity, execute step 5.4.8; if there is no available capacity, execute step 5.4.9;

[0047] Step 5.4.8, keep the outgoing edge in the reservoir pool and go to step 6;

[0048] Step 5.4.9: The total number of edges currently arriving at the reservoir pool is r, and the capacity of the reservoir pool is m. Then the edges in the reservoir pool are randomly replaced with a probability of p = r / m.

[0049] Step 6: The aggregation node detects abnormal users based on the real-time estimation results of the working nodes and the weight information table;

[0050] Step 6.1: Summarize the real-time estimation results from the working nodes to obtain the global and local triangle counts. The specific summary formula is:

[0051]

[0052] Among them, i represents the serial number of the working node, represents the global triangle count estimation result of the i-th worker node, represents the local triangle count estimation result of vertex u of the i-th working node, |W| represents the total number of working nodes, represents an estimate of the global triangle count, Represents an estimate of the local triangle count associated with vertex u.

[0053] Step 6.2: Calculate the local clustering coefficient and global clustering coefficient of all vertices to obtain a table of abnormal vertices with high clustering coefficients. The specific calculation formula is:

[0054]

[0055] Among them, d v represents the degree of vertex v in the graph, V represents all the vertices in the graph, d u represents the degree of vertex u in the graph, C represents the global clustering coefficient, C u represents the local clustering coefficient associated with vertex u;

[0056] Step 6.3, obtain abnormal users by comparing the weight information table and the high clustering coefficient abnormal vertex table;

[0057] Specifically, by comparing the weight information table with the high clustering coefficient abnormal vertex table, the vertices that exist in the high clustering coefficient abnormal vertex table but do not exist in the weight information table are abnormal users;

[0058] The beneficial effects of adopting the above technical solution are:

[0059] The present invention provides a method for detecting abnormal users in social networks based on graph flow triangle counting. Compared with the prior art, the technology proposed in the present invention achieves significant optimization in two key aspects. On the one hand, by improving the algorithm design and data processing strategy, the present invention effectively improves the estimation accuracy of triangle counting and enhances the sensitivity to changes in local network structure, thereby more accurately capturing the behavioral characteristics of abnormal users; on the other hand, by optimizing the distributed system architecture and communication protocol, the present invention greatly reduces the communication overhead between distributed nodes, thereby reducing overall resource consumption and improving execution efficiency. This dual improvement not only improves the performance of the algorithm, but also further expands its applicability in practical applications, especially in complex distributed environments with high throughput and low latency requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 This is an overall flow chart of the abnormal user detection method for social network of the present invention;

[0061] Figure 2 An example diagram of global triangles and local triangles in the graph stream of the present invention;

[0062] Figure 3It is a framework diagram of the distributed system of the present invention;

[0063] Figure 4 Flow chart of sampling edge flow in the present invention. DETAILED DESCRIPTION

[0064] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0065] The embodiment of the present invention proposes a method for detecting abnormal users in social networks based on graph stream triangle counting, innovatively designs a real-time sampling algorithm, and supports real-time triangle counting of dynamic graph streams using multiple machines in a distributed environment. Through efficient parallel computing and dynamic analysis of clustering coefficients, abnormal users in social networks are accurately identified, and the real-time and scalability of anomaly detection are achieved. The method is especially optimized for the following practical scenarios: distributed resource isolation, that is, the data storage on each working node is independent of each other, and the data on a single working node cannot be accessed or shared by other nodes, thereby ensuring the independence and security of data in the distributed system; Exactly Once traversal principle, between the master node, the working node and the aggregation node, the edge operations in the graph stream are strictly processed in chronological order, and each node can only access the edges currently being processed and the sample graph data stored locally, and cannot backtrack the edge operations that have been processed. The above design fully meets the characteristics of distributed graph stream computing, effectively reduces resource consumption and ensures processing efficiency. The overall flow chart of the method is as follows: Figure 1 As shown, the specific steps include:

[0066] like Figure 1 As shown, the following steps are included:

[0067] Step 1: Obtain a graph stream dataset that reflects the relationships between users in a real-world social network;

[0068] In the embodiment of the present invention, the Soc-Youtube-Growth dataset provided by the well-known open dataset storage platform NetworkRepository is selected. This dataset describes the relationship between all users in the online social network YouTube, contains about 3.22 million vertices, 9.37 million edges, and 12.32 million triangles, and has significant large-scale and complex characteristics.

[0069] Step 2: Preprocess the graph stream data set and randomly select a portion of the data as the training set;

[0070] The preprocessing is specifically as follows: in order to make the data set more consistent with the characteristics of the graph stream, each piece of data in the graph stream data set needs to be converted into a triple form containing two vertices and time information, that is, edge stream e = (u, v, t), where u and v are vertices, t is a timestamp, vertices represent users in a social network, and edges represent network interactions between users; then, after the current timestamp of the graph stream data set exceeds 50% of the total time, abnormal vertex information is randomly inserted to simulate the interaction behavior of abnormal users in the network. The network interaction characteristics of these abnormal users are significantly different from those of normal users, so that the performance of the abnormal user detection algorithm can be effectively tested;

[0071] In addition, for different data sets, the number of selected training sets is different, and the real-time estimation accuracy of the algorithm obtained will be different. If too few training sets are selected, the weight information table obtained through training may not cover all heavy edges; if too many training sets are selected, the weight information table obtained through training may cover unnecessary light edges. Therefore, the size of the selected training set is usually 10% to 30% of the data set.

[0072] For example, Figure 2 As shown, vertex a can form 6 local triangles, namely Δ=(a,b,c), Δ=(a,d,e), Δ=(a,d,g), Δ=(a,f,g), Δ=(a,e,f), Δ=(a,e,g), and the vertex degree is 6.

[0073] In the embodiment of the present invention, 100 vertices in the time period after the current timestamp exceeds 50% of the total time are randomly selected, and 1000 edges are inserted for each vertex, which are connected to other random vertices in an abnormal way. The inserted edges simulate the network interaction characteristics of abnormal users, such as over-connection, isolated connection or random interaction, aiming to significantly deviate from the behavior pattern of normal users; then, the first 15% of the data set is selected as the training set, which contains 1.4 million edges.

[0074] Step 3: Calculate the triangle count and vertex degree according to the training set, combine them to get the vertex weight value, and then intercept the first n% of the vertices with the largest weight as the weight information table. The weight value calculation formula is:

[0075] W e =λ 1 ×η+λ 2 (δ 1 +δ 2 ) (1)

[0076] where η represents the local triangle count of edge e = (u, v), δ 1 and δ 2 represents the degree of the edge e = (u, v) between two vertices u and v, λ 1 and λ 2represents the weight coefficient;

[0077] In the embodiment of the present invention, the weight coefficients are set to λ 1 = 0.4 and λ 2 =0.6, get the weight value of the fixed point, and intercept the first 100,000 vertices and their weight values ​​as the weight information table.

[0078] The setting of weight coefficients will affect the prediction accuracy of heavy edges. In the field of triangle counting, the importance of a vertex is measured by the number of triangles that can be formed. The more triangles there are, the more important the vertex is at the current moment. On the other hand, the degree of the vertex is considered. The more degrees of the vertex, the more likely it is to form triangles with other vertices in the future, and the more important the vertex is in the future. In addition, for different data sets, the number of intercepted vertices is different, which will have a certain impact on the prediction accuracy of the weight information table. If it is a sparse graph, there are more vertices and fewer edges, so many vertices need to be intercepted; if it is a dense graph, there are fewer vertices and more edges, so only a few vertices need to be intercepted.

[0079] Step 4: Distribute the training set from the master node to the worker nodes through a distributed architecture;

[0080] Step 4.1, set the hash function according to the number of working nodes, and calculate the hash values ​​of the two vertices u and v of the edge flow;

[0081] Step 4.2: Add a tag to the edge flow according to the hash value of u and v and the load of the working node;

[0082] Step 4-3, send the marked edge flow to the working node;

[0083] Among them, worker nodes and aggregation nodes are key components in the distributed computing architecture. In order to effectively process large-scale graph stream data, such as Figure 3 As shown, the distributed architecture specifically adopts the Master-Worker-Aggregator three-layer architecture; the master node Master is used to allocate tasks and coordinate the computing work of each worker node; the worker node Worker is used to count global and local triangles and process graph stream data locally; the aggregation node Aggregator is used to summarize the counting results of each worker node;

[0084] The distributed architecture adopts unicast and broadcast communication forms, and the hash function distributes the edge flow in combination with the load situation; if the load situation of the working node does not exceed the set threshold, the distribution situation is determined according to the vertex hash value; if the load situation of the working node exceeds the threshold, the vertex is preferentially distributed to the working node with the smallest load;

[0085] In the embodiment of the present invention, 12 virtual machines are used to simulate the distributed scenario, of which 1 is the master node, 1 is the aggregation node, and the remaining 10 are working nodes. For the 9.37 million edges of Soc-Youtube-Growth, the total number of edges received by each working node is approximately between 6.47 million and 6.64 million, and the communication load is approximately 69% to 71% of that of a single machine.

[0086] Step 5: Real-time estimation of global triangle and local triangle counts through working nodes;

[0087] Step 5.1, initialize the working node, set the sampling set to empty, and set the counter to 0;

[0088] The sampling set is composed of a waiting queue, a weighted queue and a reservoir pool in sequence, and the proportion of the three types of memory occupied is controlled by parameters, and the parameters specifically include the total capacity k of the sampling set, the waiting queue ratio α and the weighted queue ratio β; the waiting queue capacity is W=αk, the weighted queue capacity is H=(1-α)k, and the reservoir pool capacity is S=(1-α)(1-β)k;

[0089] Step 5.2, receiving edge streams and updating real-time estimates of global triangle and local triangle counts;

[0090] Among them, the real-time estimate of the current triangle count needs to be calculated based on the edge flow. The specific situation includes:

[0091] (1) If both edge u and edge v are in the reservoir, the sampling probability of the triangle is:

[0092]

[0093] (2) If one of the edges u or v is in the reservoir pool and the other is in the waiting queue or weight queue, the sampling probability of the triangle is:

[0094]

[0095] (3) If both edge u and edge v are not in the reservoir pool, the sampling probability of the triangle is:

[0096] p uvw =1 (4)

[0097] Among them, l is the total number of edges in the current graph flow.

[0098] Step 5.3: Determine whether to perform sampling according to the edge flow marking situation;

[0099] By judging whether the mark of the edge flow is equal to the sequence number of the working node, if so, the working node needs to perform sampling and execute step 5.4; if not, the edge flow is discarded and step 6 is executed;

[0100] Step 5.4: Sample the edge flow according to the weight information table, update the sampling set, and sample the sampling set according to the preset parameters and capacity, such as Figure 4 As shown, the specific steps are as follows:

[0101] Step 5.4.1: The waiting queue is a first-in-first-out queue. After receiving the edge flow, the available capacity is determined. If there is available capacity, step 5.4.2 is executed; if there is no available capacity, step 5.4.3 is executed.

[0102] Step 5.4.2, queue the edge flow and execute step 6;

[0103] Step 5.4.3, add the edge into the queue, dequeue the edge at the end of the queue, and send the dequeued edge to the weighted queue, and execute step 5.4.4;

[0104] Step 5.4.4: The weighted queue is a priority queue. After receiving the outgoing edge, a weight value is assigned to the outgoing edge according to the weight information prediction table to determine the available capacity. If there is available capacity, execute step 5.4.5; if there is no available capacity, execute step 5.4.6;

[0105] Step 5.4.5, queue the outgoing edge and execute step 6;

[0106] Step 5.4.6, queue the outgoing edge, dequeue the edge at the end of the queue, and send the outgoing edge to the reservoir pool, and execute step 5.4.7;

[0107] Step 5.4.7, the reservoir pool adopts equal probability sampling. After receiving the outbound edge, it determines the available capacity. If there is available capacity, execute step 5.4.8; if there is no available capacity, execute step 5.4.9;

[0108] Step 5.4.8, keep the outgoing edge in the reservoir pool and go to step 6;

[0109] Step 5.4.9: The total number of edges currently arriving at the reservoir pool is r, and the capacity of the reservoir pool is m. Then the edges in the reservoir pool are randomly replaced with a probability of p = r / m.

[0110] Among them, the waiting queue dynamically adjusts the sampling strategy to give priority to sampling the most active edges in the current graph flow, thereby effectively utilizing the temporal locality of the graph flow; the weight queue gives priority to retaining edges with high weight values, and assigns weight values ​​to the out-of-queue edges through the weight information prediction table. If the currently processed edge flow is in the prediction table, the edge flow is assigned a corresponding weight value based on the information in the prediction table; if the current edge flow is not in the prediction table, the edge flow is assigned a weight value of 0; the reservoir pool ensures that important data flow information is not lost under memory constraints through reservoir sampling, making the sample distribution more uniform.

[0111] In the embodiment of the present invention, the preset parameters of the sampling set are set to α=0.1 and β=0.05. For the 9.37 million edges of Soc-Youtube-Growth, the sampling set size is set to 10% of the data set, that is, k=937659, corresponding to the waiting queue capacity W=93765, the weight queue capacity H=42195 and the reservoir pool capacity S=801669. Each working node receives about 6.47 million to 6.64 million edges, of which about 1.4 million to 1.7 million edges need to be sampled, accounting for 14.9% to 18.1% of the total number of edges in the data set.

[0112] Step 6: The aggregation node detects abnormal users based on the real-time estimation results of the working nodes and the weight information table;

[0113] Step 6.1: Summarize the real-time estimation results from the working nodes to obtain the global and local triangle counts. The specific summary formula is:

[0114]

[0115] Among them, i represents the serial number of the working node, represents the global triangle count estimation result of the i-th worker node, represents the local triangle count estimation result of vertex u of the i-th working node, |W| represents the total number of working nodes, represents an estimate of the global triangle count, Represents an estimate of the local triangle count associated with vertex u.

[0116] Step 6.2: Calculate the local clustering coefficient and global clustering coefficient of all vertices to obtain a table of abnormal vertices with high clustering coefficients. The specific calculation formula is:

[0117]

[0118] Among them, d v represents the degree of vertex v in the graph, V represents all the vertices in the graph, d u represents the degree of vertex u in the graph, C represents the global clustering coefficient, C u represents the local clustering coefficient associated with vertex u;

[0119] Step 6.3, obtain abnormal users by comparing the weight information table and the high clustering coefficient abnormal vertex table;

[0120] Specifically, by comparing the weight information table with the high clustering coefficient abnormal vertex table, the vertices that exist in the high clustering coefficient abnormal vertex table but do not exist in the weight information table are abnormal users;

[0121] In the absence of abnormal users, the vertices in the two tables should be highly overlapped. Therefore, we only need to filter out the vertices that do not exist in the weight information table but are newly added in the high clustering coefficient abnormal vertex table. These vertices are abnormal users.

[0122] In the embodiment of the present invention, the aggregation node is responsible for receiving the real-time estimation values ​​calculated by the 10 working nodes, and summarizing them to obtain a global clustering coefficient of 0.07615; when generating the weight information table, the first 100,000 vertices with the highest weight values ​​are selected, and for the convenience of comparison, the high clustering coefficient abnormal vertex table also intercepts the first 100,000 vertices with the highest local clustering coefficients; after obtaining the screening results, they are compared with the preset 1000 vertices, and the accuracy rate is 89.1%. This result shows that by calculating the triangle count in the graph stream data in real time, abnormal users can be effectively identified, and it has high accuracy in large-scale social networks.

[0123] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) to form a technical solution.

Claims

1. A method for detecting abnormal users in social networks based on graph flow triangle counting, characterized in that: The following steps are involved: Step 1: Obtain a graph stream dataset that reflects the relationships between users in a real-world social network; Step 2: Preprocess the graph stream data set and randomly select a portion of the data as the training set; Step 3: Calculate the triangle count and vertex degree according to the training set, combine them to get the vertex weight value, and then intercept the first n% of the vertices with the largest weights as the weight information table; Step 4: Distribute the training set from the master node to the worker nodes through a distributed architecture; Step 5: Real-time estimation of global triangle and local triangle counts through working nodes; Step 6: The aggregation node detects abnormal users based on the real-time estimation results of the working nodes and the weight information table.

2. According to the method for detecting abnormal users in social networks based on graph flow triangle counting as described in claim 1, it is characterized in that: The preprocessing described in step 2 is specifically as follows: convert each data in the graph stream data set into a triple form containing two vertices and time information, that is, edge stream e = (u, v, t), where u and v are vertices, t is a timestamp, vertices represent users in a social network, and edges represent network interactions between users; then, after the current timestamp of the graph stream data set exceeds 50% of the total time, randomly insert abnormal vertex information to simulate the interaction behavior of abnormal users in the network.

3. According to the method for detecting abnormal users in social networks based on graph flow triangle counting as described in claim 1, it is characterized in that: The weight value calculation formula in step 3 is: W e =λ1×η+λ2(δ1+δ2) (1) Among them, η represents the local triangle count of edge e=(u,v), δ1 and δ2 represent the degrees of the two vertices u and v of edge e=(u,v), and λ1 and λ2 represent weight coefficients.

4. According to the method for detecting abnormal users in social networks based on graph flow triangle counting as described in claim 1, it is characterized in that: Step 4 specifically includes the following steps: Step 4.1, set the hash function according to the number of working nodes, and calculate the hash values ​​of the two vertices u and v of the edge flow; Step 4.2: Add a tag to the edge flow according to the hash value of u and v and the load of the working node; Step 4-3, send the marked edge flow to the working node; The distributed architecture specifically adopts the three-layer architecture of Master-Worker-Aggregator; the master node Master is used to allocate tasks and coordinate the computing work of each worker node; the worker node Worker is used to count global and local triangles and process graph stream data locally; the aggregation node Aggregator is used to summarize the counting results of each worker node; The distributed architecture adopts unicast and broadcast communication forms, and the hash function distributes the edge flow in combination with the load situation; if the load situation of the working node does not exceed the set threshold, the distribution situation is determined according to the vertex hash value; if the load situation of the working node exceeds the threshold, the vertex is preferentially distributed to the working node with the smallest load.

5. According to the method for detecting abnormal users in social networks based on graph flow triangle counting as described in claim 1, it is characterized in that: Step 5 specifically includes the following steps: Step 5.1, initialize the working node, set the sampling set to empty, and set the counter to 0; The sampling set is composed of a waiting queue, a weighted queue and a reservoir pool in sequence, and the proportion of the three types of memory occupied is controlled by parameters, and the parameters specifically include the total capacity k of the sampling set, the waiting queue ratio α and the weighted queue ratio β; the waiting queue capacity is W=αk, the weighted queue capacity is H=(1-α)k, and the reservoir pool capacity is S=(1-α)(1-β)k; Step 5.2, receiving edge streams and updating real-time estimates of global triangle and local triangle counts; Among them, the real-time estimate of the current triangle count needs to be calculated based on the edge flow. The specific situation includes: (1) If both edge u and edge v are in the reservoir, the sampling probability of the triangle is: (2) If one of the edges u or v is in the reservoir pool and the other is in the waiting queue or weight queue, the sampling probability of the triangle is: (3) If both edge u and edge v are not in the reservoir pool, the sampling probability of the triangle is: p uvw =1 (4) Among them, l is the total number of edges in the current graph flow; Step 5.3: Determine whether to perform sampling according to the edge flow marking situation; By judging whether the mark of the edge flow is equal to the sequence number of the working node, if so, the working node needs to perform sampling and execute step 5.4; if not, the edge flow is discarded and step 6 is executed; Step 5.4: Sample the edge flow according to the weight information table, update the sampling set, and sample the sampling set according to the preset parameters and capacity.

6. According to the method for detecting abnormal users in social networks based on graph flow triangle counting as described in claim 5, it is characterized in that: Step 5.4 specifically includes the following steps: Step 5.4.1: The waiting queue is a first-in-first-out queue. After receiving the edge flow, the available capacity is determined. If there is available capacity, step 5.4.2 is executed; if there is no available capacity, step 5.4.3 is executed. Step 5.4.2, queue the edge flow and execute step 6; Step 5.4.3, add the edge into the queue, dequeue the edge at the end of the queue, and send the dequeued edge to the weighted queue, and execute step 5.4.4; Step 5.4.4: The weighted queue is a priority queue. After receiving the outgoing edge, a weight value is assigned to the outgoing edge according to the weight information prediction table to determine the available capacity. If there is available capacity, execute step 5.4.5; if there is no available capacity, execute step 5.4.6; Step 5.4.5, queue the outgoing edge and execute step 6; Step 5.4.6, queue the outgoing edge, dequeue the edge at the end of the queue, and send the outgoing edge to the reservoir pool, and execute step 5.4.7; Step 5.4.7, the reservoir pool adopts equal probability sampling. After receiving the outbound edge, it determines the available capacity. If there is available capacity, execute step 5.4.8; if there is no available capacity, execute step 5.4.9; Step 5.4.8, keep the outgoing edge in the reservoir pool and go to step 6; Step 5.4.9: The total number of edges currently arriving at the reservoir pool is r, and the capacity of the reservoir pool is m. Then the edges in the reservoir pool are randomly replaced with a probability of p = r / m.

7. According to the method for detecting abnormal users in social networks based on graph flow triangle counting as described in claim 1, it is characterized in that: Step 6 specifically includes the following steps: Step 6.1: Summarize the real-time estimation results from the working nodes to obtain the global and local triangle counts. The specific summary formula is: Among them, i represents the serial number of the working node, represents the global triangle count estimation result of the i-th worker node, represents the local triangle count estimation result of vertex u of the i-th working node, |W| represents the total number of working nodes, represents an estimate of the global triangle count, represents the estimated value of the local triangle count associated with vertex u; Step 6.2: Calculate the local clustering coefficient and global clustering coefficient of all vertices to obtain a table of abnormal vertices with high clustering coefficients. The specific calculation formula is: Among them, d v represents the degree of vertex v in the graph, V represents all the vertices in the graph, d u represents the degree of vertex u in the graph, C represents the global clustering coefficient, C u represents the local clustering coefficient associated with vertex u; Step 6.3, obtain abnormal users by comparing the weight information table and the high clustering coefficient abnormal vertex table; Specifically, by comparing the weight information table and the high clustering coefficient abnormal vertex table, the vertices that exist in the high clustering coefficient abnormal vertex table but not in the weight information table are abnormal users.

Citation Information

Patent Citations

  • A method and apparatus for triangle counting in online adaptive sampling space of large-scale flow graphs

    CN113448732B

  • Social network abnormal behavior detection method and device and computer equipment

    CN116226550A

Cited By

  • Network information detection method and system based on dynamic graph flow triangle counting, terminal and storage medium

    CN121000617A