A method, device and computer device for detecting abnormal behaviors in a social network
By using auxiliary data structures and hash functions to store persistent subgraphs and non-persistent subgraph patterns separately in social networks, the problems of inefficiency and high cost in the prior art are solved, and efficient abnormal behavior detection is achieved.
Patent Information
- Application Number
- CN202310264542.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-03-17
AI Technical Summary
Existing persistent subgraph discovery technologies are inefficient in detecting abnormal behaviors of social networks and are costly to high computing and storage, making them unable to effectively mine persistent patterns.
Using an auxiliary data structure, the new k-edge subgraph is mapped into the counting slot and the bucket through a hash function, and the persistent subgraph and non-persistent subgraph mode are stored separately. Only the persistent accumulated values of the non-persistent subgraph are recorded in the counting slot, and the persistent subgraph mode is updated in real time at each timestamp.
It improves the efficiency of abnormal behavior detection, reduces memory and computing costs, and ensures the reasonable allocation of memory and the accuracy of detection.
Smart Images

Figure CN116226550B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data mining technologies, and particularly to a method, apparatus, and computer device for detecting abnormal behaviors in a social network. Background Art
[0002] Graph stream analysis is becoming increasingly important in various fields because many practical graph applications are inherently dynamic. In the past, the subgraph discovery problem of graph streams mainly focused on features such as frequency and burstiness. Persistence, as a new feature, is receiving increasing attention. Persistent subgraph discovery highlights the behavior of subgraphs recurring in many time windows, which is crucial for many practical applications (such as anomaly detection). Although persistent subgraph discovery has many interesting applications in real life, there is no ready-made solution to effectively mine persistent patterns.
[0003] One recent development is the proliferation of high-throughput, dynamically graph-structured data organized in the form of graph streams. For example, consider the knowledge graph DBpedia, which is updated daily according to the change log stream in Wikipedia. Graph stream analysis is becoming increasingly important in various fields such as subgraph matching, frequent pattern mining, and burst pattern mining. In addition to the above features, another important feature - persistence - is also receiving increasing attention. Given a subgraph pattern P and a graph stream with T tumbling windows, the persistence of P is defined as the number of time windows in which P appears. If the persistence of P is greater than a user-defined threshold, then P is said to be a persistent pattern. Persistent patterns usually indicate the occurrence of abnormal or notable events. Next, an example of detecting abnormal behaviors in a computer network is used to illustrate its basic idea.
[0004] Abnormal behaviors have Pattern 1. Security analysts can identify abnormal behaviors by monitoring the occurrence of abnormal subgraph patterns in network traffic (based on the semantics of subgraph isomorphism). As Figure 1 shown, some abnormal behaviors attempt to hide by spreading their communications across multiple time windows. As a result, these patterns cannot be detected by finding frequent subgraph patterns. To detect such threats, we should use persistence instead of frequency as an indicator. Figure 1 Shows two communication patterns and their matching results within the corresponding time windows. P1 is a pattern detected by finding a frequent subgraph pattern, which is just a general broadcast mechanism and does not provide valuable information. P2 is a pattern detected by using persistence and represents an attack pattern. P2 describes an information leak where the attacked host receives commands from a bot program and exchanges data with a compromised website that causes data leakage.
[0005] Formally, given a graph stream G, a persistence threshold δ, and an integer k, the continuous persistent pattern discovery problem is to find k-edge subgraph patterns that appear in at least δ flip windows. Although important, the continuous persistent pattern discovery problem lacks dedicated techniques for handling. A simple approach is to enumerate all possible k-edge subgraphs in each time window and then calculate the corresponding patterns of these subgraphs to verify the existence of each pattern in the current window. This method requires calculating and storing all k-edge subgraphs of each time window. In addition, it is necessary to re-execute subgraph isomorphism calculations to verify the existence of each k-edge pattern in each window, which consumes a large amount of time and memory. Therefore, advanced techniques are needed to effectively discover persistent patterns of event behaviors in order to detect abnormal behaviors in social networks in a timely and accurate manner. Summary of the Invention
[0006] Based on this, it is necessary to provide a social network abnormal behavior detection method, device, and computer device that can ensure detection efficiency and accuracy for the above technical problems.
[0007] A social network abnormal behavior detection method, the method includes:
[0008] Obtain a social network snapshot graph of the current timestamp, and extract a new set of k-edge subgraphs containing the newly inserted edges of the current timestamp from the social network snapshot graph; the new set of k-edge subgraphs includes multiple new k-edge subgraphs; the social network snapshot graph is a derived graph that includes all edges within the historical time window, all edges of the historical timestamps within the current time window, and the newly inserted edges of the current timestamp; each edge is formed by connecting 2 vertices, where the vertices represent users and the edges represent event behaviors formed by interactions between users;
[0009] Obtain the auxiliary data structure of the current timestamp; the auxiliary data structure consists of l counting slots and l buckets; each bucket corresponds to a counting slot and consists of w key-value pairs; the key in each key-value pair corresponds to a k-edge subgraph pattern, and the value corresponds to the persistent cumulative value of the k-edge subgraph pattern; the counting slot is used to record the persistent cumulative value of non-persistent subgraphs; within a time window, a key-value pair participates in persistent value counting at most once; each k-edge subgraph pattern corresponds to an event behavior;
[0010] Obtain a pre-constructed hash function, and use the hash function to map each new k-edge subgraph to the corresponding bucket. When the new k-edge subgraph is not isomorphic to the k-edge subgraph patterns of all key-value pairs in the corresponding bucket, use the hash function to map each new k-edge subgraph to the corresponding counting slot. If the minimum persistent cumulative value in the bucket corresponding to the new k-edge subgraph is less than the persistent cumulative value of the non-persistent subgraph in the corresponding counting slot, then exchange the persistent cumulative value of the counting slot and the minimum persistent cumulative value and update the key corresponding to the minimum persistent cumulative value to the pattern of the k-edge subgraph;
[0011] If there is a persistent cumulative value exceeding a preset threshold after the current time window, it is determined that the event behavior corresponding to the corresponding k-edge subgraph pattern is an abnormal behavior.
[0012] A social network abnormal behavior detection device, the device includes:
[0013] A new k-edge subgraph set extraction module, used to obtain a social network snapshot graph of the current timestamp, and extract a new k-edge subgraph set containing newly inserted edges of the current timestamp from the social network snapshot graph; the new k-edge subgraph set includes multiple new k-edge subgraphs; the social network snapshot graph is an export graph containing all edges within the historical time window, all edges of the historical timestamps within the current time window, and newly inserted edges of the current timestamp; each edge is formed by connecting 2 vertices, the vertices represent users, and the edges represent event behaviors formed by the interactions between users.
[0014] An auxiliary data structure acquisition module, used to obtain the auxiliary data structure of the current timestamp; the auxiliary data structure consists of l counting slots and l storage buckets; each storage bucket corresponds to a counting slot and consists of w key-value pairs; the key in each key-value pair corresponds to a k-edge subgraph pattern, and the value corresponds to the persistent cumulative value of the k-edge subgraph pattern; the counting slot is used to record the persistent cumulative value of the non-persistent subgraph; within a time window, a key-value pair participates in persistent value counting at most once; each k-edge subgraph pattern corresponds to an event behavior.
[0015] A persistent subgraph pattern update module, used to obtain a pre-constructed hash function, and use the hash function to map each new k-edge subgraph to the corresponding storage bucket. When the new k-edge subgraph is not isomorphic to the k-edge subgraph patterns of all key-value pairs in the corresponding storage bucket, use the hash function to map each new k-edge subgraph to the corresponding counting slot. If the minimum persistent cumulative value in the storage bucket corresponding to the new k-edge subgraph is less than the persistent cumulative value of the non-persistent subgraph in the corresponding counting slot, then exchange the persistent cumulative value of the counting slot and the minimum persistent cumulative value, and update the key corresponding to the minimum persistent cumulative value to the pattern of the k-edge subgraph.
[0016] An abnormal behavior determination module, used to determine that the event behavior corresponding to the corresponding k-edge subgraph pattern is an abnormal behavior if there is a persistent cumulative value exceeding a preset threshold after the current time window.
[0017] A computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0018] Obtain a snapshot graph of the social network at the current timestamp, and extract a new set of k-edge subgraphs containing the newly inserted edges at the current timestamp from the snapshot graph of the social network; the new set of k-edge subgraphs includes multiple new k-edge subgraphs; the snapshot graph of the social network is an exported graph containing all the edges within the historical time window, as well as all the edges with historical timestamps within the current time window and the newly inserted edges at the current timestamp; each edge is formed by connecting two vertices, where the vertices represent users and the edges represent the event behaviors formed by the interactions between users.
[0019] Obtain an auxiliary data structure at the current timestamp; the auxiliary data structure consists of l counting slots and l buckets; each bucket corresponds to a counting slot and consists of w key-value pairs; the key in each key-value pair corresponds to a k-edge subgraph pattern, and the value corresponds to the persistent cumulative value of the k-edge subgraph pattern; the counting slot is used to record the persistent cumulative value of the non-persistent subgraph; within a time window, a key-value pair participates in persistent value counting at most once; each k-edge subgraph pattern corresponds to an event behavior.
[0020] Obtain a pre-constructed hash function, and use the hash function to map each new k-edge subgraph to the corresponding bucket. When the new k-edge subgraph is not isomorphic to the k-edge subgraph patterns of all the key-value pairs in the corresponding bucket, use the hash function to map each new k-edge subgraph to the corresponding counting slot. If the minimum persistent cumulative value in the bucket corresponding to the new k-edge subgraph is less than the persistent cumulative value of the non-persistent subgraph in the corresponding counting slot, then exchange the persistent cumulative value of the counting slot and the minimum persistent cumulative value and update the key corresponding to the minimum persistent cumulative value to the pattern of the k-edge subgraph.
[0021] If there is a persistent cumulative value exceeding the preset threshold after the current time window, determine that the event behavior corresponding to the k-edge subgraph pattern is an abnormal behavior.
[0022] The above-mentioned method, device, and computer device for detecting abnormal behaviors in a social network store persistent subgraph patterns and non-persistent subgraph patterns separately, where the counting slot only stores the persistent cumulative value of the non-persistent subgraph without dividing the non-persistent subgraph pattern, and the persistent subgraph pattern is divided and persistent value counting is performed in the bucket. In this way, the memory is reasonably allocated, greatly improving the detection efficiency and reducing the memory and computing costs. At the same time, the persistent subgraph pattern will be updated in real time at each timestamp to ensure maintaining the state of reasonable memory allocation. In summary, this method can improve the detection efficiency of abnormal behaviors and reduce the detection computing costs. Description of the Drawings
[0023] Figure 1 For two communication modes and their matching results within the corresponding time window;
[0024] Figure 2 For the flow diagram of a method for detecting abnormal behaviors in a social network in an embodiment;
[0025] Figure 3 Schematic diagram of the graph stream G in an embodiment;
[0026] Figure 4 Example diagram of the auxiliary data structure in an embodiment;
[0027] Figure 5 Algorithm flow of the social network abnormal behavior detection method;
[0028] Figure 6 Example diagram of the hashing process of the auxiliary data structure in an embodiment;
[0029] Figure 7 Algorithm flow of findPP;
[0030] Figure 8 Algorithm flow of fastPP;
[0031] Figure 9 Comparison diagram of the experimental results of F1 score and throughput; among them, Figure 9 (1) Comparison diagram of the F1 score results on three datasets with default parameters, Figure 9 (2) Comparison diagram of the throughput results on three datasets with default parameters; Figure 9 (3) Comparison diagram of the F1 score results on different memories; Figure 9 (4) Comparison diagram of the throughput results on different memories;
[0032] Figure 10 Test result diagram of the influence of different parameters on recall rate, accuracy rate and throughput; among them, Figure 10 (1)-(3) are the test result diagrams of the influence of the number of key-value pairs w on recall rate, accuracy rate and throughput in sequence; Figure 10 (4)-(6) are the test result diagrams of the influence of the number of counting slots / buckets l on recall rate, accuracy rate and throughput in sequence; Figure 10 (7)-(9) are the test result diagrams of the influence of the persistence threshold δ on recall rate, accuracy rate and throughput in sequence;
[0033] Figure 11 Example diagram of the persistent pattern obtained by adopting this solution;
[0034] Figure 12 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0035] To make the objectives, technical solutions and advantages of this application more clear and understandable, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not intended to limit this application.
[0036] With the development of the Internet, Internet applications have achieved rapid development, and social media has also developed rapidly. With the development of technology, topic hype and the like have become tools for making huge profits. Topic hype is to hype a certain topic by forwarding information to each other, so as to achieve purposes such as obtaining influence in public opinion, publicity and promotion, etc. Graphs have become a common data applied to many sciences and engineering. A graph can be represented as such a structure, that is, a graph G = (V, E) is a pair of sets: a set of vertices V represents entities and a set of edges E represents the relationships or connections between entities. In computer science, a network includes nodes and edges; while in social science, the corresponding terms are actors and relationships, and these two terms have the same meaning in this invention. If the vertices in the graph are used to represent the people participating in the activity and the edges are used to represent the messages or the associations between people. Then when a media hype is initiated, in a specific time or specific scenario, multiple k-edge subgraphs are generated among the people participating in the activity. The mutual following relationships between users constitute a social network graph. Monitoring the persistence of k-edge subgraphs in the social network graph according to its dynamic changes helps to timely detect the occurrence of abnormal events in the social network and make corresponding countermeasures in a timely manner.
[0037] In one embodiment, as Figure 2 shown, a method for detecting abnormal behaviors in a social network is provided, including the following steps:
[0038] Step 202, obtain a snapshot graph of the social network at the current timestamp, and extract a new set of k-edge subgraphs containing the newly inserted edges at the current timestamp from the snapshot graph of the social network.
[0039] Among them, the new set of k-edge subgraphs includes multiple new k-edge subgraphs. The snapshot graph of the social network is a derived graph containing all the edges within the historical time window, all the edges with historical timestamps within the current time window, and the newly inserted edges at the current timestamp. Each edge is formed by connecting 2 vertices, where the vertices represent users and the edges represent the event behaviors formed by the interactions between users.
[0040] Given a graph stream G and a positive integer k, a k-edge subgraph refers to an induced subgraph that exactly has k edges existing in the graph stream G. The graph stream G refers to an ever-growing sequence of directed edges {σ1, σ2, …, σ n}, where each represents the arrival time from vertex to is t(σ iThe directed edges of (), the superscript of the vertex is the introduced vertex ID, which is used to distinguish two vertices with the same label. It is worth noting that the throughput of the graph stream is constantly changing. For simplicity of representation, only graphs with vertex labels are considered.
[0041] A time window is a series of fixed, non - overlapping, and continuous time intervals. The time window (denoted as W i ) is a time interval with a fixed duration τ in G, where the time intervals do not overlap. Specifically, the time window W i is a set of edges whose timestamps are within [t0+(i - 1)*τ, t0+i*τ), where t0 is the start time and i > 0.
[0042] A schematic diagram of the graph stream G is as Figure 3 shown. Specifically, for the edge σ2 in G, it shows that σ2 has two vertices b2 and c3, where "b" and "c" are vertex labels and the superscript is the vertex ID. For example: Xiaohong (b 1 ) is a student (b), Xiaoming (b 2 ) is a student (b), and the label of the student (b) is Xiaohong and Xiaoming. To distinguish Xiaohong and Xiaoming, vertex IDs are introduced for differentiation. In addition, the timestamp of σ2 is shown below. The graph stream is divided into three time windows, starting from the start timestamp t0 = 0, each time window has a size τ = 3, and they do not overlap.
[0043] The snapshot graph at timestamp t, denoted as G t , is the graph derived from all the edges in W i observed before time t (including time t), where t ∈ W i .
[0044] For any t ∈ W i , at time t + 1, a newly inserted edge e is obtained and added to G t to obtain G t+1 . For each newly inserted edge e in G t+1 , the symbol E k (e) is used to represent the set of new k - edge subgraphs in G t+1 that contain e. In addition, we use G i to represent the snapshot graph of the time window W i , where G i is the graph derived from all the edges in W i .
[0045] If the subgraph g k =(V g , E g ) is derived from k edges in G t , then it is called a k - edge subgraph. Define P k as Gi The set of all induced subgraphs in with k edges.
[0046] Step 204, obtain the auxiliary data structure of the current timestamp.
[0047] The auxiliary data structure consists of l counting slots and l buckets. Each bucket corresponds to a counting slot and consists of w key-value pairs. The key in each key-value pair corresponds to a k-edge subgraph pattern, and the value corresponds to the persistent cumulative value of the k-edge subgraph pattern. Within a time window, a key-value pair participates in persistent value counting at most once. The counting slot is used to record the persistent cumulative value of non-persistent subgraphs; each k-edge subgraph pattern corresponds to an event behavior.
[0048] The semantic isomorphism relationship divides the subgraph set Pk into m equivalent classes, denoted by Each equivalent class is called a subgraph pattern. Note that can be obtained by deleting the ID (or timestamp) of the vertex (or edge) corresponding to the k-edge subgraph in . For simplicity, use the abbreviation P i to represent the general pattern We define the frequency fre(P i in each time window as i G i ) as the number of k-edge subgraphs in Use the symbol PS to represent a set of different k-edge patterns in G. Each item in PS is a binary tuple (P, per(P)), where P is a k-edge pattern and per(P) is the persistence value of pattern P.
[0049] Persistence refers to a specific pattern of occurrence behavior based on the number of windows in which a k-edge subgraph pattern appears in the graph stream. The persistence measure of pattern P is to calculate the number of time windows in which P appears. Note that we do not consider that a persistent pattern must appear in all time windows, because this is a very special case. Therefore, the persistence measure of a persistent pattern should exceed the user-defined threshold δ. The formal definition of a persistent pattern is as follows: Given a graph stream G, a k-edge pattern P, and a persistence threshold δ. If per(P) ≥ δ, then P is a persistent pattern.
[0050] Step 206: Obtain a pre-constructed hash function. Use the hash function to map each new k-edge subgraph to a corresponding bucket. When the new k-edge subgraph is not isomorphic to the k-edge subgraph patterns of all key-value pairs in the corresponding bucket, use the hash function to map each new k-edge subgraph to a corresponding counting slot. If the minimum persistent cumulative value in the bucket corresponding to the new k-edge subgraph is less than the persistent cumulative value of the non-persistent subgraph in the corresponding counting slot, then swap the persistent cumulative value of the counting slot and the minimum persistent cumulative value and update the key corresponding to the minimum persistent cumulative value to the pattern of the k-edge subgraph.
[0051] As Figure 4 shown, an example graph of the auxiliary data structure is provided. In Figure 4 , the first part is an array, which is associated with the hash function h1(·) and maps each subgraph to a slot in the array. Each slot is a counter. The array has l counting slots counter, where the i-th counting slot in the array is denoted as C1[i]; the second part consists of l buckets, where the i-th bucket is denoted as B[i]. Each bucket corresponds to a counting slot in the array, that is, B[i] corresponds to C1[i]. Each bucket has ω key-value pairs. If P is stored in B[i], then B[h1(P)][P] represents the corresponding key-value pair of the k-edge pattern P, where the key of B[i] is P and the value is the corresponding counter. Initially, all status fields are True and the value of each counter is 0. Since within a time window, a key-value pair participates in persistent value counting at most once, a counting status field True or False can be set in the key-value pair to indicate whether the key-value pair and / or the counting slot has participated in persistent value counting within the current time window: when the key-value pair and / or the counting slot has not participated in persistent value counting within the current time window, the corresponding counting status field is True; when the key-value pair and / or the counting slot has participated in persistent value counting once within the current time window, the corresponding counting status field is False. It should be noted that the status field can be designed according to requirements, and only an exemplary description is given here.
[0052] When the new k-edge subgraph is not isomorphic to the k-edge subgraph patterns of all key-value pairs in the corresponding bucket, it indicates that the new k-edge subgraph may be a non-persistent subgraph. Therefore, it is mapped to the counting slot for recording the persistent cumulative value of the non-persistent subgraph. Further, if the minimum persistent cumulative value in the bucket corresponding to the new k-edge subgraph is less than the persistent cumulative value of the non-persistent subgraph in the corresponding counting slot, it means that the pattern corresponding to the new k-edge subgraph has higher persistence than the pattern corresponding to the minimum persistent cumulative value. Therefore, the persistent cumulative value in the counting slot and the minimum persistent cumulative value are exchanged, and the key corresponding to the minimum persistent cumulative value is updated to the pattern of the k-edge subgraph. At the same time, it should be noted that the status field of the counting slot for the exchange turns to False.
[0053] It can be seen that the auxiliary data structure of this solution ensures that each newly generated k-edge subgraph only needs to access one array and one slot, greatly improving the time efficiency. And to detect more persistent k-edge patterns, the key idea of this solution is to separate persistent and non-persistent k-edge patterns. This is because considering that if non-persistent patterns occupy too many buckets, then there will not be enough space for persistent patterns to be stored, resulting in a reduction in recall rate when memory is limited.
[0054] Step 208, if there is a persistent cumulative value exceeding the preset threshold after the current time window, determine that the event behavior corresponding to the k-edge subgraph pattern is an abnormal behavior.
[0055] In the above method for detecting abnormal behaviors in a social network, the persistent subgraph patterns and non-persistent subgraph patterns are stored separately. Among them, the counting slot only stores the persistent cumulative value of the non-persistent subgraph without dividing the non-persistent subgraph patterns. The division of persistent subgraph patterns and the counting of persistent values are only carried out in the bucket. In this way, the memory is reasonably allocated, greatly improving the detection efficiency and reducing the memory and calculation costs. At the same time, the persistent subgraph patterns will be updated in real time at each timestamp to ensure maintaining the state of reasonable memory allocation. In summary, this can improve the detection efficiency of events of interest while reducing the detection calculation costs.
[0056] It should be understood that although Figure 2 the steps in the flowchart of Figure 2At least some of the steps may include multiple sub-steps or multiple stages, and these sub-steps or stages do not necessarily need to be executed and completed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages does not necessarily need to be sequential, but can be executed alternately or in turns with at least some of the sub-steps or stages of other steps or other steps.
[0057] In one embodiment, when the new k-edge subgraph is isomorphic to the k-edge subgraph pattern of the key-value pair in the corresponding bucket and the key-value pair has not participated in the persistent value counting within the current time window, the key-value pair participates in the persistent value counting and updates the persistent cumulative value of the corresponding k-edge subgraph pattern.
[0058] In one embodiment, if the minimum persistent cumulative value in the bucket corresponding to the new k-edge subgraph is not less than the persistent cumulative value of the non-persistent subgraph in the corresponding counting slot and the counting slot has not participated in the persistent value counting within the current time window, the counting slot participates in the persistent value counting and updates the persistent cumulative value of the corresponding non-persistent subgraph.
[0059] As Figure 5 shown, an algorithmic process for providing a social network abnormal behavior detection method is provided. That is, for each subgraph g in E k (e), there are the following two cases: k
[0060] 1) g k is isomorphic to one of the key-value pairs in B[h1(g k )]. If the status of the counter in B[h1(g k )][j] is True, check whether g k is isomorphic to B[h1(g k )][j]. If isomorphic, increment the counter by 1 in B[h1(g k )][j] and set the status of the counter to False;
[0061] 2) g k is not isomorphic to any key-value pair in B[h1(g k )] and C1[h1(g k )].state == True. Increment the counter in B[h1(g k ) and set the status of the counter to False. Then find the minimum counter in B[h1(g k )] to compare with the counter in C1[h1(g k )] in order to determine whether the corresponding pattern P is persistent enough to be stored in the bucket B[h1(g k )]. Use B[h1(g k )] minto store the key-value pair with the minimum persistence in B[h1(g k )]. If C1[h1(g k )].value > B[h1(g k )] min .value, it means that the estimated persistence of P is greater, and P should be stored in B[h1(g k )]. Therefore, swap C1[h1(g k )] and B[h1(g k )][B[h1(g k )] min .key].value, and swap P and B[h1(g k )][B[h1(g k )] min .key].key to obtain the updated auxiliary data structure; if C1[h1(g k )].value ≤ B[h1(g k )] min .value, then do not swap.
[0062] As Figure 6 shown, assume ω = 2, and provide an example diagram of the hashing process for the auxiliary data structure. In the current time window, it is found that the subgraph is hashed to B[h1(g k )][1]. Since is isomorphic to P5, and the status of the counter in B[h1(g k )] is True, TFD + increments the counter by 1 and changes the status to False; is not isomorphic to any of the key-value pairs in, so is mapped to the counting slot The counting slot is updated to (False, 4) because 4 is greater than the minimum counter 3 in the bucket, so the key is set to P6, which is the pattern of the subgraph , and (False, 4) and (True, 3) are swapped. After that, in order to map the subgraph Because is not isomorphic to any of the key-value pairs in, TFD+ maps to the counting slot, and the counting slot is updated to (False, 4).
[0063] In one embodiment, the step of using a hash function to map each new k-edge subgraph to the corresponding bucket or counting slot includes:
[0064] Encode each new k-edge subgraph into a string representation using graph invariants so that isomorphic subgraphs are mapped to corresponding buckets or counting slots, specifically including:
[0065] Concatenate the degrees and labels of each vertex of each new k-edge subgraph e = (v i , v j , t(e)) together as the new label l(v) of the corresponding vertex; where v i , v j are the vertices in the new k-edge subgraph e, and t(e) is the edge formed by the corresponding vertices in the new k-edge subgraph e;
[0066] Obtain the new label l(e) = (l(v i ), l(v j )) of each edge in the new k-edge subgraph according to the new labels of the vertices;
[0067] Assign a weight w(e) to each edge according to the order in which the single-edge pattern corresponding to the edge appears in the snapshot graph; the single-edge pattern is the subgraph pattern when k = 1. Among them, the earlier the single-edge pattern appears for the first time, the smaller the corresponding weight. For example: if the single-edge pattern A -> B is the first edge to appear in the graph stream, then the weights of all edges isomorphic to the single-edge pattern A -> B are all 1; then the single-edge pattern B -> C appears, and the weight is 2; and so on. The purpose of doing this is to ensure that isomorphic subgraphs have the same encoding, that is, graph invariants, so that the encodings of two isomorphic subgraphs are the same.
[0068] If w(e i ) < w(e j ), then e i < e j ; that is, the single-edge pattern corresponding to e i appears earlier than the single-edge pattern corresponding to e j ;
[0069] If w(e i ) = w(e j ) ∪ l(e i ) < l(e j ), then e i < e j , where l(e i ) < l(e j ) means that the vertex degree of e i is smaller; the vertex degree is equal to the number of edges that the vertex participates in forming;
[0070] If w(e i ) > w(e j ), then e i > e j ;
[0071] Obtain the corresponding encoded string representation {l(e1), …, l(e n )} according to the weight of each edge of the new k-edge subgraph, where e i <e i+1 .
[0072] In one embodiment, the step of determining whether the new k-edge subgraph is isomorphic to the k-edge subgraph pattern of the key-value pair in the bucket or the non-persistent subgraph in the counting slot includes:
[0073] Obtain the new k-edge subgraph and the k-edge subgraph pattern of the key-value pair in the bucket or the non-persistent subgraph in the counting slot where, represents the vertices in the new k-edge subgraph, represents the edges formed by the corresponding vertices in the new k-edge subgraph; represents the vertices of the k-edge subgraph pattern of the key-value pair, represents the edges formed by the corresponding vertices of the k-edge subgraph pattern of the key-value pair;
[0074] When there exists a bijective function f(·) from to and satisfies 1) and 2) , the new k-edge subgraph is isomorphic to the k-edge subgraph pattern of the key-value pair in the bucket or the non-persistent subgraph in the counting slot, otherwise it is not isomorphic. The function L(·) is used to maintain the labels of the vertices.
[0075] Next, experimental verification is carried out for this solution to prove its performance. All algorithms are implemented in C++ and run on a PC with an Intel i7 3.50GHz CPU and 32GB of memory. In all experiments, space is reserved to store the entire graph stream. Each quantitative test is repeated 5 times.
[0076] 1. Datasets. For each dataset, it is divided into 50 time windows, i.e., T = 50.
[0077] · Enron is an email communication network consisting of 86K entities (such as employees) and 297K edges (such as emails), and the timestamps correspond to the communication data.
[0078] · Offshore contains a total of 839K offshore entities (such as companies), 3.6 million relationships (such as establishment) and 433 labels, covering offshore entities and financial activities, and the timestamps correspond to the active days.
[0079] · Facebook is a social network consisting of 415K entities (such as users) and 2.1 million edges (such as messages), and the timestamps correspond to the release dates.
[0080] 2. Solutions for comparison. The following three solutions to the persistent subgraph pattern discovery problem are compared.
[0081] · findPP: A baseline method for mining persistent patterns;
[0082] · fastPP: An advanced algorithm framework using the data - assisted structure TFD;
[0083] · fastPP+: This solution improves fastPP by combining the optimized TFD+.
[0084] The introductions of findPP and fastPP are given respectively as follows:
[0085] 1) findPP: A straightforward way to discover persistent patterns on a graph stream is to enumerate all possible k - edge subgraphs when a new time window arrives, and then partition the set of k - edge subgraphs into different equivalence classes to verify the occurrence of each k - edge pattern. If the persistence measure of a k - edge pattern exceeds the user - defined threshold, it is returned as a persistent pattern. More details are described below. As Figure 7 shown, the algorithm flow of findPP is provided. The set PS is used to store different k - edge patterns in G. Each item in PS is a pair (P, per(P)), where P is a k - edge pattern and per(P) is the persistence value of pattern P. Whenever a new window W i appears, findPP updates PS by calling computePer (lines 2 - 3). Then, for each pattern P in PS, findPP verifies whether the persistence value of P satisfies the persistence threshold δ (lines 4 - 5). Finally, it returns all persistent patterns (line 6).
[0086] Function computePer. computePer first calls findSubgraph(·) to compute the set of k - edge subgraphs P in Gi k (line 1). Specifically, whenever an edge insertion e occurs at timestamp t (t ∈ W i ), findSubgraph explores a candidate subgraph space in a tree - like manner in G t to compute E k (e). Each node represents a candidate subgraph, where a child node extends from its parent node by one edge. To avoid duplicate enumeration of subgraphs, findSubgraph checks whether two subgraphs are composed of the same edges at each layer in the tree space. After processing all edge insertions in W i , the set of k - edge subgraphs P in G i can be obtained. k. To compute the corresponding k-edge patterns, computePer calls evaluateFre(·) to partition the subgraphs in P k into equivalence classes according to subgraph isomorphism calculation. Each equivalence class can represent a pattern P (line 2). If fre(P, G i ) ≥ 1, computePer further checks whether P ∈ PS through subgraph isomorphism calculation; if so, set per(P) ← per(P) + 1, otherwise, it adds (P, per(P) = 1) to the set PS (lines 3 - 6). Among them, the tree space is an auxiliary data structure for gradually expanding a newly inserted edge into a k-edge subgraph. Being in the same layer means that the number of edges of the k-edge subgraph is the same.
[0087] There are three main steps in findPP. (1) During the k-edge subgraph enumeration process, given an inserted edge e in G t , let n be the average number of vertices of the subgraph with a radius of k expanded from e. findSubgraph spends to explore all k-edge subgraphs containing e. (2) During the PS update process, let σ be the average unit time for checking whether two k-edge subgraphs are isomorphic. evaluateFre partitions the set of k-edge subgraphs into m equivalence classes in O(N·(N 2 -1)·σ) time. Let M be the number of patterns in PS. computerper takes O(m·M·σ) to update PS. (3) findPP takes O(1) to return the persistent patterns.
[0088] 2) fastPP. First, analyze the shortcomings of the baseline algorithm findPP, and then design a new auxiliary data structure TFD to effectively mine the persistent patterns on the graph stream, which can significantly reduce the memory cost and computational cost. Why is the cost high? The scalability of the findPP algorithm is insufficient to handle large graph streams. First, to find the k-edge patterns in each time window, findPP needs to compute and store all k-edge subgraphs in each time window, which consumes a large amount of time and memory. Second, during the PS update process, findPP needs to re-execute the subgraph isomorphism calculation for each pattern in the current window to check whether it exists in PS.
[0089] As Figure 8As shown, the algorithmic process of fastPP is provided. First, it calls initializePer to initialize the TFD (line 1). Then, it updates the TFD by calling updateTFD to calculate the persistence of each pattern P in the TFD when a new time window arrives (lines 2 - 3). After processing the current time window, it sets the state of the counter in the TFD to True (lines 4 - 5). Then, for each non - empty bucket B i [j] in the TFD, it checks whether the value of B i [j].value meets the persistence threshold δ (lines 6 - 7). Finally, it returns all persistent patterns (line 8).
[0090] Function updateTFD. updateTFD processes the snapshot graphs in W in ascending order (line 1). i Whenever an edge insertion e occurs at timestamp t (t ∈ W i ), updateTFD calls findSubgraph to calculate E k (e) (lines 2 - 3). For each subgraph g k ∈ E k (e), the TFD first selects a hash function h1(·) to map g k to a bucket B i in the array B i [h i (g k )], and then checks whether the state of B i [h i (g k )].state is True to avoid over - estimation (lines 4 - 6). If so, there are two cases: (1) g k is isomorphic to the pattern P in the bucket B i [h i (g k ). updateTFD increments the persistence of P by 1 and sets the state of B i [h i (g k )].state to B i [h i (gk)].state (lines 7 - 9). (2) gk is not isomorphic to the pattern P and the bucket is empty. updateTFD first calculates the pattern of g k by removing the vertex IDs and edge timestamps of g k , and then inserts (P, (state = False, per(P) = 1)) into the bucket (lines 10 - 12). Finally, updateTFD returns the updated TFD (line 13).
[0091] Compared with the baseline solution, fastPP directly maps each newly generated k-edge subgraph into a fixed bucket in the TFD to calculate the persistence of the pattern, rather than calculating and storing all k-edge subgraphs in each time window, which significantly reduces the memory cost and time cost. More importantly, once the state of the counter indicates that the corresponding bucket has been counted in the current time window, fastPP will not need to verify the existence of the pattern in the bucket in the window, which avoids duplicate subgraph matching calculations.
[0092] Algorithm fastPP + Compared with the algorithm fastPP, the process of updating the auxiliary data structure is different.
[0093] 3. Metrics. The following four metrics are used:
[0094] · Recall Rate (RR): The ratio of the number of reported persistent patterns to the number of persistent patterns.
[0095] · Precision Rate (PR): The ratio of the number of reported persistent patterns to the number of reported patterns.
[0096] · F1 Score: 2×RR×PR / (RR + PR). It is calculated based on the accuracy and recall rate of the test, which is also a measure of the test accuracy.
[0097] · Throughput: The number of edge increment updates that can be processed per second (KIPS).
[0098] 4. Parameter Settings. There are four parameters: the number d of hash functions (in the TFD) and the number ω of key-value pairs in the bucket (in the TFD + ), the number l of buckets (in the TFD) and the number l of cells in the bucket (TFD + ), and the persistence threshold δ.
[0099] Specifically, d is changed from 4 to 16, with a default value of 12, ω is changed from 4 to 16, with a default value of 12, and l is changed from 4 to 32, with a default value of 16. δ can be set by domain scientists according to domain knowledge, selected from 10 to 30, with a default value of 20. In addition, the subgraph size k = 4 is determined. If not specified otherwise, when changing a certain parameter, the values of other parameters will be set to their default values.
[0100] 5. Performance Evaluation
[0101] EXP-1: Impact of the dataset. First, evaluate the performance of findPP, fastPP, and fastPP+ on Enron, Offshore, and Facebook using a limited amount of memory (50MB, 250MB, and 150MB respectively).
[0102] F1 score. As Figure 9 (1) shows, the results show the F1 scores of fastPP + and its competitors on three datasets with default parameters. Similar results can also be observed under other parameter settings. As Figure 9 (1) shows, it can be observed that the F1 score of fastPP + is higher than that of all other competitors, and fastPP is also higher than findPP. For example, on Facebook, the F1 score of fastPP+ is higher than 95%, and that of fastPP + > 90%, while that of fastPP + is lower than 80%. The main reason is that findPP needs sufficient memory to store all k-edge subgraphs to ensure accuracy, which will lead to poor performance when the memory is limited. Compared with findPP, our data structure does not need to store any subgraphs, which is less affected by the memory size. The reason why fastPP + is inferior to fastPP is that some non-persistent patterns may be first mapped into the TFD, resulting in insufficient positions for subsequent persistent patterns when the number of positions in the TFD is limited. As a result, some persistent patterns are misjudged as non-persistent patterns. Note that the F1 score of findPP can reach 100% under sufficient memory because findPP accurately calculates the persistence of each pattern under sufficient memory.
[0103] Throughput. From Figure 9 (2), it is also observed that the throughput of fastPP + is always higher than that of other algorithms, and fastPP is also higher than findPP. Specifically, fastPP + performs up to 1.4 times better than fastPP on Enron, while fastPP performs up to 4.7 times better than findPP on Offshore. This is because findPP first needs to partition P k and calculate k-edge subgraph patterns based on subgraph isomorphism. Then, findPP needs to re-execute subgraph isomorphism calculation for each pattern to check whether it exists in PS, which also results in a high computational cost. In contrast, fastPP directly maps it into fixed buckets in the TFD to form corresponding k-edge patterns, which can avoid expensive calculations. In addition, fastPP + can further improve efficiency because it ensures that each subgraph only needs to access one counter and one bucket. fastPP + performs slightly differently in the three datasets, but the trends are very similar. The results show the robustness of fastPP + , so in the following experiments, only the Offshore dataset is used.
[0104] EXP-2: Impact of Memory. Evaluated fastPP + and its competitors for accuracy and speed with different memory sizes on Offshore. In the experiment, the memory size was set between 200MB and 400MB.
[0105] F1 score. In Figure 9 (3), it can be seen that the increase in memory size has a great impact on findPP because there is not enough space to store all subgraphs to accurately calculate the persistence of each pattern. It was also found that the increase in memory size has little impact on fastPP and fastPP + and the least impact on fastPP + . This is because they use auxiliary data structures without storing all k-edge subgraphs. Because TFD and TFD + only store k-edge patterns and their persistence. Moreover, fastPP + guarantees that non-persistent patterns will be quickly replaced and will not occupy too much space. Therefore, our algorithm runs well with very limited memory.
[0106] Throughput. As shown in Figure 9 (4), as expected, fastPP + is much faster than other algorithms. We can see that the increase in memory size reduces the throughput of findPP and has little impact on the throughput of fastPP and fastPP + . The reason is that findPP takes less time to partition P k into m equivalent classes. The throughput of findPP is much lower because findPP causes redundant subgraph matching calculations. Due to fewer memory accesses, the throughput of fastPP + is the best.
[0107] EXP-3: Impact of Parameters. We evaluated the RR, PR, and throughput of fastPP and fastPP + using different parameters on Offshore with a fixed memory size (i.e., 250MB). Note that when changing parameters, we kept other parameters as default values. The results for other datasets were consistent.
[0108] Impact of d / ω. As shown in Figure 10(As shown in (1)-(3), in this experiment, we varied d from 4 to 16. In particular, we observed that an increase in d can increase the RR of fastPP and decrease its throughput, and the PR of fastPP is always 1. When ω increases, the throughput generally decreases, and the RR and PR generally increase. The reason is that for larger d or ω, there are more opportunities for potential persistent patterns to be stored in the TFD and TFD + and the RR of fastPP and fastPP + will increase. Moreover, for larger ω, there are more positions to store persistent patterns, which can protect more potential persistent patterns from replacement and hash conflicts, and the PR of fastPP + will also increase. Since the TFD precisely calculates the persistence of patterns, the PR of fastPP is always 1. However, the throughput of fastPP will decrease because it has to check d - 1 buckets for each edge insertion. And the throughput of fastPP + will be reduced because when ω is larger, the number of memory accesses is less. Therefore, choosing the appropriate d or ω is a trade-off between accuracy and throughput. The larger d or ω, the higher the accuracy, while the lower the throughput. If the application requires high throughput, we should reduce d or ω. If the application requires high accuracy, we should increase d or ω.
[0109] The influence of l. As Figure 10 (shown in (4)-(6), the experimental results show that an increase in l can increase the RR of fastPP and fastPP + and decrease its throughput. This is because when l increases, there are more positions in the TFD and TFD + and we can detect more patterns simultaneously. However, since we need to calculate the persistence of patterns in each position of the TFD and TFD + in each time window, it will result in more subgraph matching computations. l does not affect the PR of fastPP and fastPP + because l does not affect the persistence of k-edge patterns in the TFD and TFD + .
[0110] The influence of δ. As Figure 10 (shown in (7)-(9), the experimental results show that an increase in δ can increase the RR of fastPP and fastPP + . This is because for smaller δ, the baseline results may be huge and we can only detect a fixed number of patterns in the TFD and TFD + . Therefore, the RR is lower. We also found that an increase in δ can increase the fastPP +PRs can safely filter more false positives due to persistence constraints. fastPP and fastPP + The throughput of is insensitive to δ because δ does not affect TFD and TFD + The number of k-edge patterns in
[0111] We demonstrated the practicality of the persistent pattern discovery problem through a case study on Offshore. Figure 11 Shows the persistent patterns obtained by our solution. It reveals a significant shift of active bearer companies from the "British Virgin Islands" to "Panama". This pattern can help us detect abnormal behaviors of these active bearer companies and analyze the reasons for the abnormal behaviors. Through analysis, we know that the reason for the abnormal behavior is that the British Virgin Islands has cracked down on bearer shares. For the British Virgin Islands, it is necessary to guard against the risk of capital outflows in order to better establish an offshore financial market.
[0112] In one embodiment, a social network abnormal behavior detection device is provided, including:
[0113] A new k-edge subgraph set extraction module, configured to obtain a social network snapshot graph at the current timestamp, and extract a new k-edge subgraph set containing newly inserted edges at the current timestamp from the social network snapshot graph; the new k-edge subgraph set includes multiple new k-edge subgraphs; the social network snapshot graph is a derived graph containing all edges within a historical time window, all edges with historical timestamps within the current time window, and newly inserted edges at the current timestamp; each edge is formed by connecting 2 vertices, the vertices represent users, and the edges represent event behaviors formed by interactions between users;
[0114] An auxiliary data structure acquisition module, configured to acquire an auxiliary data structure at the current timestamp; the auxiliary data structure consists of l counting slots and l buckets; each bucket corresponds to a counting slot and consists of w key-value pairs; the key in each key-value pair corresponds to a k-edge subgraph pattern, and the value corresponds to the persistent cumulative value of the k-edge subgraph pattern; the counting slot is used to record the persistent cumulative value of non-persistent subgraphs; within a time window, a key-value pair participates in persistent value counting at most once; each k-edge subgraph pattern corresponds to an event behavior;
[0115] The persistent subgraph pattern update module is used to obtain a pre - constructed hash function, map each new k - edge subgraph to the corresponding bucket by using the hash function. When the new k - edge subgraph is not isomorphic to the k - edge subgraph patterns of all key - value pairs in the corresponding bucket, map each new k - edge subgraph to the corresponding counting slot by using the hash function. If the minimum persistent cumulative value in the bucket corresponding to the new k - edge subgraph is less than the persistent cumulative value of the non - persistent subgraph in the corresponding counting slot, exchange the persistent cumulative value of the counting slot and the minimum persistent cumulative value and update the key corresponding to the minimum persistent cumulative value to the pattern of the k - edge subgraph;
[0116] The abnormal behavior determination module is used to determine that the event behavior corresponding to the k - edge subgraph pattern is an abnormal behavior if there is a persistent cumulative value exceeding a preset threshold after passing through the current time window.
[0117] For the specific limitations of the social network abnormal behavior detection device, reference can be made to the limitations of the social network abnormal behavior detection method in the above text, which will not be elaborated here. Each module in the above - mentioned social network abnormal behavior detection device can be implemented in whole or in part by software, hardware, and their combination. The above - mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above - mentioned modules.
[0118] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 12 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non - volatile storage medium and an internal memory. The non - volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non - volatile storage medium. The database of the computer device is used to store graph stream data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a social network abnormal behavior detection method.
[0119] Those skilled in the art can understand that Figure 12 the structure shown in
[0120] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method in the above embodiment are implemented.
[0121] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method in the above embodiment are implemented.
[0122] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0123] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0124] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A method for detecting abnormal behaviors in a social network, characterized in that The method includes: Obtaining a snapshot graph of a social network at the current timestamp, and extracting a new set of k-edge subgraphs containing the newly inserted edges at the current timestamp from the snapshot graph of the social network; the new set of k-edge subgraphs includes multiple new k-edge subgraphs; the snapshot graph of the social network is a derived graph containing all edges within a historical time window, all edges with historical timestamps within the current time window, and newly inserted edges at the current timestamp; each edge is formed by connecting two vertices, where the vertices represent users and the edges represent event behaviors formed by interactions between users. Obtaining an auxiliary data structure at the current timestamp; the auxiliary data structure consists of l counting slots and l buckets; each bucket corresponds to a counting slot and consists of w key-value pairs; the key in each key-value pair corresponds to a k-edge subgraph pattern, and the value corresponds to the persistent cumulative value of the k-edge subgraph pattern; the counting slots are used to record the persistent cumulative values of non-persistent subgraphs; within a time window, a key-value pair participates in persistent value counting at most once; each k-edge subgraph pattern corresponds to an event behavior. Obtaining a pre-constructed hash function, and using the hash function to map each new k-edge subgraph to the corresponding bucket. When the new k-edge subgraph is not isomorphic to the k-edge subgraph patterns of all key-value pairs in the corresponding bucket, use the hash function to map each new k-edge subgraph to the corresponding counting slot. If the minimum persistent cumulative value in the bucket corresponding to the new k-edge subgraph is less than the persistent cumulative value of the non-persistent subgraph in the corresponding counting slot, then exchange the persistent cumulative value of the counting slot and the minimum persistent cumulative value, and update the key corresponding to the minimum persistent cumulative value to the pattern of the k-edge subgraph. If there is a persistent cumulative value exceeding a preset threshold after the current time window, determine that the event behavior corresponding to the k-edge subgraph pattern is an abnormal behavior.
2. The method according to claim 1, wherein The key-value pairs and counting slots also include a counting status field; the counting status field is True or False. When the key-value pair and / or counting slot has not participated in persistent value counting within the current time window, the corresponding counting status field is True. When the key-value pair and / or counting slot has participated in persistent value counting once within the current time window, the corresponding counting status field is False.
3. The method according to claim 1, wherein The method further includes: When the new k-edge subgraph is isomorphic to the k-edge subgraph pattern of the key-value pair in the corresponding bucket and the key-value pair has not participated in persistent value counting within the current time window, the key-value pair participates in persistent value counting and updates the persistent cumulative value of the corresponding k-edge subgraph pattern.
4. The method according to claim 1, wherein The method further includes: If the minimum persistent cumulative value in the bucket corresponding to the new k-edge subgraph is not less than the persistent cumulative value of the non-persistent subgraph in the corresponding counting slot and the counting slot has not participated in persistent value counting within the current time window, the counting slot participates in persistent value counting and updates the persistent cumulative value of the corresponding non-persistent subgraph.
5. The method according to any one of claims 1 to 4, characterized in that, The step of using the hash function to map each new k-edge subgraph to the corresponding bucket or counting slot includes: Encoding each new k-edge subgraph into a string representation using graph invariants such that isomorphic subgraphs are mapped to corresponding buckets or counting slots, specifically including: For each new k-edge subgraph e = (v i , v j , t(e)), concatenate the degrees and labels of each vertex as the new label l(v) of the corresponding vertex; where v i , v j are vertices in the new k-edge subgraph e, and t(e) is the edge formed by the corresponding vertices in the new k-edge subgraph e; The new label l(e) of each edge in the new k-edge subgraph is obtained according to the new label of the vertex, where l(e) = (l(v i ), l(v j )); Assigning a weight w(e) to each edge according to the order in which the single-edge patterns corresponding to the edges of the social network snapshot graph appear; among them, the earlier the single-edge pattern appears for the first time, the smaller the corresponding weight; If w(e i ) < w(e j ), then e i < e j ; If w(e i ) = w(e j ) ∪ l(e i ) < l(e j ), then e i < e j , where l(e i ) < l(e j ) means that the vertex degree of e i is smaller; the vertex degree is equal to the number of edges formed by the vertex; If w(e i ) > w(e j ), then e i > e j ; Obtain the corresponding encoded string representation {l(e1), …, l(e n )} according to the weights of each edge of the new k-edge subgraph corresponding to the corresponding weights, where e i <e i+1 .
6. The method according to any one of claims 1 to 4, characterized in that The steps of determining whether the new k-edge subgraph is isomorphic to the k-edge subgraph pattern of the key-value pair of the bucket or the non-persistent subgraph in the counting slot include: Obtain a new k-edge subgraph and a non-persistent subgraph in the k-edge subgraph pattern or count slot of the key-value pair of the storage bucket wherein represents a vertex in the new k-edge subgraph represents an edge formed by the corresponding vertices in the new k-edge subgraph; represents a vertex of the k-edge subgraph pattern of the key-value pair represents an edge formed by the corresponding vertices of the k-edge subgraph pattern of the key-value pair; When there exists a bijective function f(·) from to and satisfies 1) and 2) the new k-edge subgraph is isomorphic to the k-edge subgraph pattern of the key-value pairs of the bucket or the non-persistent subgraph in the counting slot; otherwise, it is not isomorphic. The function L(·) is used to maintain the labels of vertices.
7. A social network abnormal behavior detection device, characterized in that, The device includes: A new k-edge subgraph set extraction module, configured to obtain a social network snapshot graph of the current timestamp, and extract a set of new k-edge subgraphs containing newly inserted edges of the current timestamp from the social network snapshot graph; the set of new k-edge subgraphs includes multiple new k-edge subgraphs; the social network snapshot graph is a derived graph including all edges within a historical time window, all edges of historical timestamps within the current time window, and newly inserted edges of the current timestamp; each edge is formed by connecting 2 vertices, the vertices represent users, and the edges represent event behaviors formed by interactions between users; An auxiliary data structure acquisition module, configured to obtain an auxiliary data structure of the current timestamp; the auxiliary data structure consists of l counting slots and l buckets; each bucket corresponds to a counting slot and consists of w key-value pairs; the key in each key-value pair corresponds to a k-edge subgraph pattern, and the value corresponds to the persistent cumulative value of the k-edge subgraph pattern; the counting slot is used to record the persistent cumulative value of the non-persistent subgraph; within a time window, a key-value pair participates in persistent value counting at most once; each k-edge subgraph pattern corresponds to an event behavior; A persistent subgraph pattern update module, configured to obtain a pre-constructed hash function, and map each new k-edge subgraph to the corresponding bucket using the hash function. When the new k-edge subgraph is not isomorphic to the k-edge subgraph patterns of all key-value pairs in the corresponding bucket, map each new k-edge subgraph to the corresponding counting slot using the hash function. If the minimum persistent cumulative value in the bucket corresponding to the new k-edge subgraph is less than the persistent cumulative value of the non-persistent subgraph in the corresponding counting slot, exchange the persistent cumulative value of the counting slot and the minimum persistent cumulative value, and update the key corresponding to the minimum persistent cumulative value to the pattern of the k-edge subgraph; An abnormal behavior determination module, configured to determine that the event behavior corresponding to the k-edge subgraph pattern is an abnormal behavior if there is a persistent cumulative value exceeding a preset threshold after the current time window.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Data structure anomaly detection method and device, storage medium and computer equipment
CN111400290A
Efficient indexed data structures for persistent memory
US20220027349A1