A social network data analysis method, system and electronic device
Through the principle of maximum breadth priority and multi-order attribute perception strategy, the problem of large communication overhead in the existing technology is solved, and efficient graph data processing and load balancing are achieved.
Patent Information
- Application Number
- CN202310239697.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-13
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-03-13
AI Technical Summary
The communication overhead of existing social network data analysis methods is large and cannot meet actual needs.
The principle of maximum breadth priority is used to access the vertices in the social network graph, and divide the vertices into the corresponding partitions according to the multi-order attribute perception strategy. Consider the number of neighbors and common neighbors of the vertices in each partition, and dynamically balance the load to reduce communication overhead.
It improves the efficiency of graph division, supports larger-scale graph data processing, reduces communication overhead, and meets actual needs.
Smart Images

Figure CN116302527B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of social network data analysis, and more specifically, relates to a social network data analysis method, system and electronic device. Background Art
[0002] In today's Internet era, graph data is a common data structure in current Internet companies. In China, there are Weibo, Tencent, Alibaba, etc., and in foreign countries, there are Twitter, Google, etc. In these companies, a huge number of users form many large social networks. These companies use graph structures to represent the relationships between users. For example, some user interactions such as likes, comments, and follows are the edges between user vertices. Through effective graph data analysis, people can understand the meaning of data more deeply. Nowadays, there are already many effective methods to extract useful information from graph data, so as to help companies better improve their products and services.
[0003] With the rapid development of the Internet, the rapidly growing scale of enterprise graph data makes the processing of social network graphs become increasingly difficult. For example, Facebook's graph data contains more than two billion user vertices and more than one trillion edges representing relationships such as follows, fans, and likes; Alibaba's graph network also contains more than one billion users and two billion commodity vertices. Generally, if the method of extracting graph data information runs on a single machine with a multi-core CPU, it can only handle graph processing with at most several million vertices, which obviously cannot meet the current needs. Therefore, the mainstream approach is to process large-scale graph data in a distributed environment, divide the graph data into multiple parts, and perform operations on multiple vertices in parallel. On the one hand, this can improve the overall computing performance; on the other hand, as long as there are enough computing nodes, it can support the processing of graph data with a scale of hundreds of millions of vertices. However, the partitioning of graph data is not as simple as imagined. Different from other linear data structures, the graph data itself also contains rich information in its structure. In most graph processing methods, to extract the information of a vertex from the graph, it is usually necessary to refer to the information of the vertex's neighbors, and even the neighbors of the neighbors. This shows that there are a large number of data dependencies inside the graph, and the partitioning of the graph needs to reduce the damage of dependencies.
[0004] While taking into account data dependencies, the performance of the distributed system also needs to be considered. More specifically, when a social network graph is divided into several subgraphs for running in a distributed system, the edge that stores two vertices on different computing nodes is called a cut edge. When operating on a cut edge, the information of both vertices is required simultaneously, which involves the problem of data synchronization. To ensure that the vertex information obtained at both ends is the latest, communication overhead is generated at this time. Therefore, the more cut edges there are, the greater the communication overhead in the distributed system. Generally speaking, communication is the main overhead in the distributed system.
[0005] In summary, the existing social network data analysis methods have a large communication overhead and cannot meet the actual needs. Summary of the Invention
[0006] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a social network data analysis method, system and electronic device, aiming to solve the problem that the existing social network data analysis methods have a large communication overhead and cannot meet the actual needs.
[0007] To achieve the above object, the present invention provides a social network data analysis method, which is characterized by including the following steps:
[0008] Store the data of the social network in the form of a social network graph, where each vertex in the social network graph represents a user. If there is a social relationship between users, there is an edge connection between the corresponding two vertices;
[0009] Determine the vertex with the largest degree in the social network graph. According to the principle of breadth-first search with the largest degree, visit the neighbor vertices of this vertex in descending order of degree. Subsequently, according to the same principle, visit the neighbor vertices of the neighbor vertices of this vertex until each vertex in the social network graph is visited;
[0010] Arrange all the visited vertices in the order of visit, and divide all the visited vertices into multiple batches according to the preset parallel granularity and evenly distribute them to each thread according to the arranged order;
[0011] In each thread, divide each vertex into the corresponding partition according to the multi-order attribute perception strategy; wherein, the multi-order attribute perception strategy considers the number of neighbors of the vertex in each partition and the number of common neighbors of the vertex and the vertices in each partition, and the multi-order attribute perception strategy has a dynamic load balancing constraint on the load of each partition to ensure the load balance of multiple partitions; the vertices with similar characteristics in the social network are divided into the same partition to reduce the communication overhead when analyzing social network data in a distributed system.
[0012] In a possible implementation manner, the vertices in the social network graph are visited according to the following steps:
[0013] (1.1) Find the vertex with the largest degree in the social network graph, visit the vertex with the largest degree and sort its neighbors in descending order of degree;
[0014] (1.2) Start visiting the neighbor vertices, and the visiting order is sorted by degree, and the neighbor with a higher degree is visited first;
[0015] (1.3) Mark the vertices that have been visited during each visit. When visiting a new vertex each time, check whether it has been visited. If it has been visited, skip this vertex;
[0016] (1.4) After all the n - order neighbors of the vertex with the maximum degree have been visited, in the order of visiting the n - order neighbors, perform steps (1.2) and (1.3) for each n - order neighbor, and perform the same visit to the (n + 1) - order neighbors of the vertex with the maximum degree until each vertex in the social network graph has been visited; where n represents the number of times steps (1.2) and (1.3) are repeatedly executed.
[0017] In a possible implementation, divide all the visited vertices into multiple batches according to the arrangement order, specifically:
[0018] Store all the visited vertices into an array according to the visit order. Determine the size b of the batch according to the preset number of threads t and the number of vertices v in the vertex stream; where,
[0019] Each thread determines its unique id number, id = 0, 1, 2, 3......, and then calculates the starting index sid = b * id and the ending index eid = (b + 1) * id of the vertices responsible for division by each thread according to b and id;
[0020] When the vertex id processed by the thread exceeds the ending index eid or exceeds the array range, stop processing.
[0021] In a possible implementation, in each thread, divide each vertex into the corresponding partition according to the multi - order attribute perception strategy, specifically:
[0022] (3.1) Each thread traverses its own vertex stream. For each vertex, then traverse all the partitions;
[0023] (3.2) Calculate the division score S(v, p) for each currently dividing vertex v and each partition p; where the division score S(v, p) consists of three parts: is the penalty coefficient, which is related to the preset hyperparameter and the number of vertices that have been divided into the partition currently. The penalty intensity changes with the current division situation to control the load balance of each partition. The value of the hyperparameter can be adjusted within a preset range to control whether to give priority to load balance or minimum cut in the division; is the first - order attribute, considering the number of neighbors of vertex v in partition p; is the second - order attribute, considering the number of common neighbors of vertex v and the vertices in partition p;
[0024] (3.3) For the vertex currently being partitioned, compare its partitioning scores in each partition, select the partition with the highest partitioning score, and partition v into the corresponding partition.
[0025] In a possible implementation, the step (3.2) specifically includes the following steps:
[0026] (3.2.1) Calculate the penalty coefficient according to the preset hyperparameter β, the number of vertices currently partitioned into the partition, the number of partitions, and the number of vertices that have been partitioned where |P| is the number of vertices in partition p, k is the number of partitions, and |V| is the number of vertices that have been partitioned, then
[0027] (3.2.2) Use a high-speed intersection algorithm to intersect the neighbor set N(v) of the current partitioning vertex v and the vertex set Q(p) that has currently been partitioned into p, obtain the result set R, and use the modulus of R as the first-order attribute value;
[0028] (3.2.2.1) Sort the vertices in N(v) and Q(p) in ascending order of vertex numbers, and compare the moduli of the two sets. Point the pointer s to the first element of the smaller set, and point the pointer l to the first element of the larger set;
[0029] (3.2.2.2) Let the element pointed to by the pointer s be S, and the element pointed to by the pointer l be L. Compare the sizes of S and L: If S is equal to L, add S to R, and both s and l point to the next element; if S is less than L, only s points to the next element; if S is greater than L, set the initial step size r equal to 1, make l point to the position of l + r, and compare S and L again. If S is still greater than L, then make r = r * 2 until S is less than L; at this time, use binary search to search for the same S' as S in the range of the set where l is located. If found, add S' to R and make l point to the next position after S'; if not found, make s point to the next element and l point to the position of l + r + 1; of the range of the set where l is located. If found, add S' to R and make l point to the next position after S'; if not found, make s point to the next element and l point to the position of l + r + 1;
[0030] (3.2.2.3) Repeat the comparison in (3.2.2.2) until s or l points to the end of the set where they are located, return the set R, and enter step (3.2.3);
[0031] (3.2.3) Traverse the set R, count the common neighbors of each vertex r in R and the current partitioning vertex v, and add up all the counted numbers of common neighbors as the second-order attribute value;
[0032] (3.2.4) Calculate the partitioning score based on the penalty coefficient, the first-order attribute, and the second-order attribute.
[0033] In a second aspect, the present invention provides a social network data analysis system, comprising:
[0034] A social network graph acquisition module, configured to store the data of the social network in the form of a social network graph, where each vertex in the social network graph represents a user, and if there is a social relationship between users, there is an edge connection between the corresponding two vertices;
[0035] A vertex access module, configured to determine the vertex with the largest degree in the social network graph, and according to the maximum degree breadth-first principle, sequentially access the neighbor vertices of the vertex in descending order of degree, and then, according to the same principle, access the neighbor vertices of the neighbor vertices of the vertex until each vertex in the social network graph is accessed;
[0036] A thread allocation module, configured to arrange all the accessed vertices in the order of access, and divide all the accessed vertices into multiple batches according to a preset parallel granularity in the arranged order, and evenly allocate them to each thread;
[0037] A vertex partitioning module, configured to partition each vertex into a corresponding partition in each thread according to a multi-order attribute perception strategy; wherein, the multi-order attribute perception strategy considers the number of neighbor vertices of the vertex in each partition and the number of common neighbor vertices of the vertex and the vertices in each partition, and the multi-order attribute perception strategy has a load dynamic balance constraint on each partition to ensure load balance of multiple partitions; vertices with similar characteristics in the social network are partitioned into the same partition to reduce communication overhead when analyzing social network data in a distributed system.
[0038] In a possible implementation manner, the vertex access module accesses vertices according to the following steps: (1.1) Find the vertex with the largest degree from the social network graph, access the vertex with the largest degree and sort its neighbors in descending order of degree; (1.2) Start accessing the neighbor vertices, and the access order is sorted according to degree, and the neighbor with a higher degree is accessed first; (1.3) Mark the vertices that have been accessed during each access, and check whether the vertex has been accessed when accessing a new vertex each time, and skip the vertex if it has been accessed; (1.4) After all the n-order neighbor vertices of the vertex with the largest degree have been accessed, perform steps (1.2) and (1.3) on each n-order neighbor according to the access order of the n-order neighbors, and perform the same access on the (n + 1)-order neighbor vertices of the vertex with the largest degree until each vertex in the social network graph is accessed; where n represents the number of times steps (1.2) and (1.3) are repeatedly executed.
[0039] In a possible implementation, the thread allocation module divides all the visited vertices into multiple batches according to the arrangement order. Specifically, it stores all the visited vertices in an array according to the visit order, and determines the size b of the batch based on the preset number of threads t and the number v of vertices in the vertex stream. Among them, Each thread determines its unique id number, id = 0, 1, 2, 3......, and then calculates the starting index sid = b * id and the ending index eid = (b + 1) * id of the vertices responsible for division by each thread according to b and id. When the vertex id processed by the thread exceeds the ending index cid or exceeds the array range, the processing stops.
[0040] In a possible implementation, in each thread, the vertex partitioning module divides each vertex into the corresponding partition according to the multi-order attribute perception strategy. Specifically: (3.1) Each thread traverses its own vertex stream, and for each vertex, traverses all the partitions again; (3.2) Calculate the partitioning score S(v, p) for each currently partitioning vertex v and each partition p. Among them, the partitioning score S(v, p) consists of three parts: is the penalty coefficient, which is related to the preset hyperparameter and the number of vertices already partitioned into the partition. The penalty intensity changes with the current partitioning situation to control the load balance of each partition. The value of the hyperparameter can be adjusted within a preset range to control whether to give priority to load balance or minimum cut in the partitioning; is the first-order attribute, considering the number of neighbors of vertex v in partition p; is the second-order attribute, considering the number of common neighbors of vertex v and the vertices in partition p; (3.3) For the currently partitioning vertex, compare its partitioning scores in each partition, select the partition with the highest partitioning score, and divide v into the corresponding partition.
[0041] In a third aspect, the present application provides an electronic device, including: at least one memory for storing a program; at least one processor for executing the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any one of the possible implementations of the first aspect.
[0042] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program. When the computer program runs on a processor, it causes the processor to execute the method described in the first aspect or any one of the possible implementations of the first aspect.
[0043] Fifth aspect, the present application provides a computer program product, which, when running on a processor, causes the processor to execute the method described in the first aspect or any possible implementation of the first aspect.
[0044] Generally speaking, compared with the prior art, the above technical solution conceived by the present invention has the following beneficial effects:
[0045] For the social network data analysis method provided by the present invention, during the partitioning process, considering that the scale of the social network graph data in the real world is too large and the efficiency of the traditional serial method of partitioning vertex by vertex is too low, the partitioning task is parallelized through reasonable parallel task partitioning. Compared with the traditional method, the graph partitioning efficiency is greatly improved, and it can support the partitioning of larger-scale graph data. In addition, the present invention also optimizes the algorithm for high-speed intersection finding for intersection, further improving the efficiency. Moreover, during the parallelization process, the present invention uses the maximum degree breadth-first algorithm to generate vertex streams. When dividing into batches, it can retain more correlations between vertices, reduce the deterioration of the partitioning effect caused by parallelization, and reduce the communication overhead in the process of processing social network graph data while maintaining the original social network data processing effect.
[0046] For the social network data analysis method provided by the present invention, by calculating the multi-order attributes and penalty parameters between vertices and partitions, it jointly determines whether a vertex is partitioned into the partition, solving the problem that only single attributes are considered in the existing methods. Compared with the strategy of only considering the number of neighbors adopted by the traditional method, the data after partitioning by the present invention can be more suitable for the current complex social network graph processing method, can effectively reduce the communication overhead, and meet the actual requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a flowchart of the social network data analysis method provided by an embodiment of the present invention;
[0048] Figure 2 is a flowchart of the parallel multi-order attribute-aware streaming graph partitioning method provided by an embodiment of the present invention;
[0049] Figure 3 is an example diagram of the parallel multi-order attribute-aware streaming graph partitioning method provided by an embodiment of the present invention;
[0050] Figure 4 is an example diagram of the fast intersection finding method provided by an embodiment of the present invention;
[0051] Figure 5 is an example diagram of the instance graph data for calculating the partitioning score provided by an embodiment of the present invention;
[0052] Figure 6It is the architecture diagram of the social network data analysis system provided by the embodiments of the present invention. Detailed implementation manners
[0053] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0054] The term "and / or" in this document is a relationship description of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The symbol " / " in this document indicates that the associated objects are in an "or" relationship. For example, A / B represents A or B.
[0055] The terms "first" and "second" in the description and claims of this document are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first response message and the second response message are used to distinguish different response messages, rather than to describe the specific order of the response messages.
[0056] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0057] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of processing units refers to two or more processing units; a plurality of elements refers to two or more elements.
[0058] Figure 1 It is the flowchart of the social network data analysis method provided by the embodiments of the present invention; as Figure 1 shown, it includes the following steps:
[0059] S101, store the data of the social network in the form of a social network graph, where each vertex in the social network graph represents a user. If there is a social relationship between users, there is an edge connection between the corresponding two vertices;
[0060] S102, determine the vertex with the largest degree in the social network graph. According to the principle of breadth-first search with the largest degree, visit the neighbor vertices of this vertex in the order of decreasing degree. Subsequently, according to the same principle, visit the neighbor vertices of the neighbors of this vertex until each vertex in the social network graph is visited;
[0061] S103. Arrange all the visited vertices in the order of visit, and divide all the visited vertices into multiple batches according to the preset parallel granularity in the arranged order, and evenly distribute them to each thread.
[0062] S104. In each thread, divide each vertex into the corresponding partition according to the multi-order attribute perception strategy; wherein, the multi-order attribute perception strategy considers the number of neighbors of the vertex in each partition and the number of common neighbors of the vertex and the vertices in each partition, and the multi-order attribute perception strategy has a dynamic load balancing constraint on the load of each partition to ensure the load balance of multiple partitions; vertices with similar characteristics in the social network are divided into the same partition to reduce the communication overhead when analyzing social network data in a distributed system.
[0063] It should be noted that graph partitioning is involved in the process of social network data processing. In graph partitioning, the optimization objectives include two items: load balance and minimum cut. Load balance is to make the performance of multiple computers in a distributed system similar; a cut includes cut edges and cut points, which refer to the edges or points that exist in two partitions at the same time. When the system processes a cut, communication overhead often occurs. The minimum cut is to reduce the communication cost between computers. Usually, blindly reducing the cut edges often leads to load imbalance and affects the overall performance. For example, in an extreme case, when all vertices are divided into the same computing vertex, the cut edges reach the minimum. Therefore, in the graph partitioning method, it is necessary to comprehensively consider both load balance and minimum cut, make a trade-off, and keep the distributed system with better load balance while keeping the minimum cut as much as possible, so as to improve the overall performance.
[0064] In addition, traditional algorithms usually process the entire graph, and the overhead for large-scale graphs is difficult to estimate. It may even require a huge amount of space just to load the entire graph into memory. A better way is to use streaming graph partitioning. For example, the LDG and Fennel algorithms no longer analyze all vertices before partitioning, but convert the vertex set into a vertex stream, and each vertex individually decides which vertex it is divided into. And when determining the partitioning, both load balance and as few cuts as possible are taken into account. However, on the one hand, the growth rate of today's graph scale is astonishing, and these algorithms are all single-threaded. When facing ultra-large graphs, the partitioning time may even offset the performance improvement brought by the partitioning; on the other hand, the methods of graph processing are becoming more and more complex, and traditional partitioning algorithms usually only consider the number of neighbors of vertices and cannot meet the increasingly complex graph processing algorithms.
[0065] Therefore, how to design a parallel graph partitioning method considering multiple metrics is an urgent problem to be solved by those skilled in the art. In order to better partition social network graph data, the embodiments of the present invention will elaborate in detail on the parallel multi-order attribute-aware streaming graph partitioning method used in the social network data processing process, specifically as Figure 2 shown, the method includes the following steps:
[0066] (1) Starting from the vertex with the maximum degree in the graph, according to the maximum degree breadth-first principle, access the next vertex from its neighbors, and continue to access using the same principle until all vertices are traversed, and sort the vertex stream in the access order.
[0067] (2) Consider the vertices as arranged horizontally in one dimension, and evenly divide the vertex stream into multiple batches according to the parallel granularity, and evenly distribute these batches to each thread.
[0068] (3) Each thread processes the above vertex stream batch by batch in the vertex order, calculates scores for each partition area for each vertex according to the multi-order attribute-aware strategy, and then partitions the vertex into the area with the highest score.
[0069] Specifically, step (1) includes:
[0070] (1.1) Find the vertex with the largest degree in the graph, access the vertex and sort its neighbors in descending order of degree.
[0071] (1.2) Start accessing its neighbors, and the access order is sorted by degree, and the neighbors with higher degrees are accessed first.
[0072] (1.3) Mark the vertices that have been accessed during each access, and check whether a vertex has been accessed when accessing a new vertex each time. If it has been accessed, skip the vertex.
[0073] (1.4) When all neighbors have been accessed, start the same access process for the neighbors of the neighbors in the access order of the previous round until each vertex has been accessed.
[0074] It can be understood that starting from the vertex with the largest degree to access the vertices in the graph, after accessing a vertex, mark it as having been accessed. And sort its neighbors in descending order of degree, add the unaccessed neighbors to the queue to be accessed. Access the vertices in the order of the queue to be accessed until all vertices have been accessed, and record the access order of the whole process.
[0075] Specifically, step (2) includes:
[0076] (2.1) Store the vertex sequence into an array in the order in (1). Determine the batch size b according to the preset number of threads t and the number of vertices v in the vertex stream. Generally, take
[0077] (2.2) Each thread obtains its own unique sequential id number, id = 0, 1, 2, 3......, and then calculates the starting index sid = b * id and the ending index eid = (b + 1) * id of the vertices responsible for partitioning by each thread according to b and id.
[0078] (2.3) When the id of the vertex processed by the thread exceeds eid or exceeds the array range, stop processing. If there are still unpartitioned vertices in the array, re-partition a batch of vertices from the unprocessed vertex area for processing until all vertices are processed.
[0079] Furthermore, each thread traverses the vertex sequence it is responsible for. For each vertex v, traverse all partitions. For each v and p, by calculating the score And for each vertex, the scores with all partitions will be calculated and saved together.
[0080] Among them, step (3) specifically includes:
[0081] (3.1) Each thread traverses its own vertex stream V t , and for each vertex, traverse all partitions again.
[0082] (3.2) Calculate the partitioning score S(v, p) for each current partitioned vertex v and each partition p. The score consists of three parts: is the penalty coefficient ω to control the load balance of each partition; is the first-order attribute, considering the number of neighbors of v in p; is the second-order attribute, considering the number of common neighbors with the vertices in p.
[0083] (3.2.1) Calculate the penalty coefficient according to the hyperparameter β is related to the degree of the current partition, including the number of vertices that have been partitioned, the number of vertices partitioned into p, etc., so as to ensure load balance. By adjusting β, it is possible to change whether to give priority to load balance or minimum cut in the partitioning.
[0084] (3.2.2) Take the intersection of the neighbor set N(v) of the current partitioned vertex v and the vertex set Q(p) that has been partitioned into p. Use a high-speed intersection algorithm during the intersection process to obtain the resulting set R, and take the modulus of R as the first-order attribute Value
[0085] (3.2.2.1) Sort the vertices of N(v) and Q(p) in ascending order of vertex numbers. And compare the sizes of the moduli of the two sets. Point the pointer s to the first element of the smaller set, and point the pointer l to the first element of the larger set.
[0086] (3.2.2.2) Set the initial step size r to 1, and compare the elements pointed to by the pointers s and l. If the element pointed to by s is equal to the element pointed to by l, add the result to the result set R, and move both the pointers s and l backward to point to the next element; if the element pointed to by l is greater than the element pointed to by s, perform a binary search for the element pointed to by s between the position pointed to by l and the previous position pointed to by l. If found, add it to the result set and point l to the next position of the search result. If not found, move s one position backward. Finally, reset r to 1; if the element pointed to by l is less than the element pointed to by s, move s backward by r positions, double r, and execute this step again.
[0087] (3.2.2.3) As the two pointers move, when s or l points to the end of the set, return the result set R and enter step (3.2.3).
[0088] (3.2.3) Traverse the set R, and count the sum of the number of common neighbors between the vertex r in R and v as the second-order attribute Value
[0089] (3.2.4) Multiply the scores of the three parts and store them in the score statistics area.
[0090] (3.3) For the vertex being partitioned currently, compare its partitioning scores in each partition, and select the partition with the highest score. Lock the data structure for statistical partitioning, add v to the partition, and then unlock it.
[0091] It should be noted that the quality of the partitioning strategy largely determines the performance of the graph stream partitioning. The main idea of the partitioning is: arrange the vertices in a certain order, and then partition the vertices into different nodes in turn until all vertices are partitioned. The present invention uses the method of calculating the "partitioning score" to determine which node the current vertex is partitioned into, that is, calculate the partitioning scores of the vertex to be partitioned and each node, and then partition the vertex into the node with the highest score. Figure 3 This is an example of vertex stream partitioning. How to calculate this proximity score is the core of the whole method, which should take into account both reducing the cut edges and keeping the node load balanced as much as possible. Obviously, the partitioning score of a vertex and a node represents the closeness of the new graph formed by the vertex and the existing vertex set in the node.
[0092] It can be understood that, in order to reduce the overhead in the process of set intersection, the present invention adopts a high-speed intersection method. The general intersection algorithm is to first sort the two sets from small to large, and then use two pointers to point to the smaller ends of the sets respectively. Compare the elements pointed to by the two pointers. If the elements are the same, add them to the result set. Otherwise, move the pointer pointing to the smaller element one position backward, and repeat the comparison until the sets are traversed. However, this intersection algorithm is only applicable to the case where the scales of the two sets are not very different. If the scales of the sets are too different, there will be a large number of useless comparisons. Through observation, it can be found that since the point stream has a gradually decreasing degree, the scale of the neighbor vertex set of a vertex changes from a large scale to a small scale. At the same time, as the partitioning progresses, the number of vertices in a node gradually increases, so the scale of this set changes from a small scale to a large scale. Therefore, most of the time, the scales of the two sets have a large gap, so there are a large number of invalid comparisons. Therefore, the present invention adopts a fast intersection algorithm to accelerate the intersection, such as Figure 4 , when moving the pointer, the fixed step size is no longer 1. When the elements pointed to by the pointer are continuously small, the initial position of the pointer is p, and the offset of each movement doubles, that is, the offset relative to the initial position is: 2 0 , 2 1 , 2 2 ,......, 2 n . When the pointed element is no longer a smaller element, use binary search in [p + 2 (n-1) , p + 2 n to find whether there are the same elements, which can avoid a large number of invalid comparisons and make the pointer pointing to the smaller element move quickly.
[0093] Such as Figure 5 shown, when calculating the proximity score of vertex e and node p: equals 2, cn(x, y) represents the number of common neighbors of x and y, then:
[0094] In the traditional graph partitioning methods LDG and Fennel, for how to calculate the tightness between a vertex and a vertex set, only the neighbor vertices of the vertex are simply used. Undoubtedly, the more neighbor vertices of the vertex to be partitioned in the vertex set, the higher the tightness between the vertex and the vertex set. However, considering that modern graph processing algorithms often use a large number of common neighbors, this second-order property, to cope with more complex scenarios. Here, the common neighbor means that there is an edge between two vertices, and another vertex has an edge with both of these two vertices, then this vertex is called the common neighbor of these two vertices. Therefore, in order to better adapt to complex graph processing algorithms, the present invention takes multi-order properties into account in the calculation of the proximity score.
[0095] Specifically, the penalty coefficient is calculated according to the hyperparameter β |P| is the number of vertices in partition p, k is the number of partitions, and |V| is the number of vertices that have been partitioned. Then
[0096] It can be understood that when a node reaches a relatively high degree of closeness, it is easier for subsequent vertices to reach a relatively high proximity score with this node, because a high degree of closeness indicates that there are more high-degree vertices in the node, and vertices and high-degree vertices are obviously more likely to have neighbor relationships and common neighbor relationships. If not restricted, all vertices will gradually be partitioned into the same node, and at this time, the distributed system will degenerate into a single machine. Therefore, the present invention penalizes the score according to the load of the current node, and the higher the load, the heavier the penalty.
[0097] The present invention calculates the load penalty coefficient using the number of vertices that have been partitioned currently instead of the total number of vertices. The penalty intensity changes with the current partitioning situation, which is called dynamic partitioning. If the total number of vertices is used to calculate the penalty coefficient, at the beginning of the partitioning, |P| is small, it will approach 1, and then there is almost no penalty effect. When a certain number of vertices are partitioned into a node, it will approach 0, and then the penalty is too heavy, and vertices are almost no longer partitioned into this node. Finally, the partitioning effect is equivalent to always partitioning into one node, stopping when the average load size is reached, and then always partitioning vertices into another node, and so on. However, using the currently partitioned vertices to calculate will make the load between nodes increase and decrease reciprocally and grow evenly at the same time.
[0098] Regarding the value of β, the strictness of load balancing. When β is smaller, it will be smaller, the load penalty will be strict, and the partitioning effect may be weakened. When β is larger, it will be larger, the load penalty will be smaller, which may lead to load imbalance.
[0099] In addition, for each vertex, add it to the vertex queue of the partition with the largest partition score, and sort it according to the vertex number size when adding. First, obtain the mutex of this partition before insertion, and then perform the insertion after obtaining the mutex. Release the mutex of this partition after the insertion is completed.
[0100] Using the above parallel multi-order attribute-aware streaming graph partitioning method, not only can large-scale graph data be partitioned efficiently, but also the multi-order attribute awareness can ensure that the partitioning result is sufficiently effective in the face of high-order graph processing methods.
[0101] The graph partitioning of the present invention is performed on multiple graph datasets of different scales. Compared with the traditional graph partitionings LDG and Fennel, it can reduce the partitioning time by several times. And it can partition a dataset with tens of millions of vertices such as Twitter at the hour level, while LDG takes several days. In addition to the ultra-high performance, the present invention also provides sufficient effectiveness. In the downstream task HuGE, an entropy-based high-order random walk algorithm, it can provide a performance improvement of up to 2 times. Therefore, the parallel multi-order attribute-aware streaming graph partitioning method provided by the present invention can ensure the effectiveness, efficiency, and scalability of graph partitioning.
[0102] Figure 6 It is the architecture diagram of the social network data analysis system provided by the embodiment of the present invention; as Figure 6 shown, it includes:
[0103] A social network graph acquisition module 610, which is used to store the data of the social network in the form of a social network graph. Each vertex in the social network graph represents a user. If there is a social relationship between users, there is an edge connection between the corresponding two vertices;
[0104] A vertex access module 620, which is used to determine the vertex with the largest degree in the social network graph. According to the maximum degree breadth-first principle, it sequentially accesses the neighbor vertices of this vertex in the order of decreasing degree. Subsequently, according to the same principle, it accesses the neighbor vertices of the neighbors of this vertex until each vertex in the social network graph is accessed;
[0105] A thread allocation module 630, which is used to arrange all the accessed vertices in the order of access, and according to the preset parallel granularity, divide all the accessed vertices into multiple batches in the arranged order and evenly distribute them to each thread;
[0106] A vertex partitioning module 640, which is used to partition each vertex into the corresponding partition in each thread according to the multi-order attribute-aware strategy; wherein, the multi-order attribute-aware strategy considers the number of neighbor vertices of the vertex in each partition and the number of common neighbor vertices of the vertex and the vertices in each partition, and the multi-order attribute-aware strategy has a dynamic load balancing constraint on the load of each partition to ensure the load balancing of multiple partitions; the vertices with similar characteristics in the social network are partitioned into the same partition to reduce the communication overhead when analyzing social network data in a distributed system.
[0107] It should be understood that the above device is used to execute the method in the above embodiment. For the corresponding program modules in the device, their implementation principles and technical effects are similar to the descriptions in the above method. The working process of this device can refer to the corresponding process in the above method, which will not be elaborated here.
[0108] Based on the method in the above embodiments, an embodiment of the present application provides an electronic device. The device may include: at least one memory for storing a program and at least one processor for executing the program stored in the memory. Wherein, when the program stored in the memory is executed, the processor is used to execute the method described in the above embodiments.
[0109] Based on the method in the above embodiments, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program runs on a processor, the processor is caused to execute the method in the above embodiments.
[0110] Based on the method in the above embodiments, an embodiment of the present application provides a computer program product, and when the computer program product runs on a processor, the processor is caused to execute the method in the above embodiments.
[0111] It can be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0112] The method steps in the embodiments of the present application may be implemented in a hardware manner, or may be implemented by a processor executing software instructions. The software instructions may be composed of corresponding software modules, and the software modules may be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may be located in an ASIC.
[0113] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
[0114] It can be understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application.
[0115] Those skilled in the art can easily understand that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for analyzing social network data, characterized in that It includes the following steps: Store the data of the social network in the form of a social network graph, where each vertex in the social network graph represents a user. If there is a social relationship between users, there is an edge connection between the corresponding two vertices; Determine the vertex with the largest degree in the social network graph. According to the principle of breadth-first search with the largest degree, visit the neighbor vertices of this vertex in descending order of degree. Subsequently, according to the same principle, visit the neighbor vertices of the neighbors of this vertex until each vertex in the social network graph is visited; Arrange all the visited vertices in the order of visit, and divide all the visited vertices into multiple batches according to the preset parallel granularity, and evenly distribute them to each thread; In each thread, divide each vertex into the corresponding partition according to the multi-order attribute perception strategy; wherein, the multi-order attribute perception strategy considers the number of neighbors of the vertex in each partition and the number of common neighbors of the vertex and the vertices in each partition, and the multi-order attribute perception strategy has a dynamic load balancing constraint on the load of each partition to ensure the load balance of multiple partitions; the vertices with similar characteristics in the social network are divided into the same partition to reduce the communication overhead when analyzing the social network data in the distributed system; In each thread, divide each vertex into the corresponding partition according to the multi-order attribute perception strategy, specifically: (3.1) Each thread traverses its own vertex stream. For each vertex, traverse all partitions again; (3.2)For each vertex currently being partitioned v and each partition p calculate the partitioning score S(v,p) ; where the partitioning score consists of three parts: ; is a penalty coefficient, which is related to the preset hyperparameter and the number of vertices already partitioned into the partition. The penalty intensity changes with the current partitioning situation to control the load balance of each partition. The value of the hyperparameter can be adjusted within a preset range to control whether to give priority to load balance or minimum cut in the partitioning; is a first-order attribute that considers the number of neighbors of vertex v in partition p ; is a second-order attribute that considers the number of common neighbors of vertex v and the vertices in partition p ; (3.3)For the vertex being partitioned currently, compare its partitioning scores in each partition, select the partition with the highest partitioning score, and v partition it into the corresponding partition.
2. The method according to claim 1, wherein Access the vertices in the social network graph according to the following steps: (1.1) Find the vertex with the largest degree from the social network graph, access the vertex with the largest degree and sort its neighbors in descending order of degree; (1.2) Start accessing the neighbor vertices, and the access order is sorted by degree, and the neighbor with a higher degree is accessed first; (1.3) Mark the vertices that have been accessed during each access. When accessing a new vertex each time, check whether it has been accessed. If it has been accessed, skip this vertex; (1.4) After all the nth-order neighbors of the vertex with the largest degree have been accessed, according to the access order of the nth-order neighbors, execute steps (1.2) and (1.3) for each nth-order neighbor, and perform the same access to the (n + 1)th-order neighbors of the vertex with the largest degree until each vertex in the social network graph is accessed; where n represents the number of times steps (1.2) and (1.3) are repeatedly executed.
3. The method according to claim 1, wherein Divide all the visited vertices into multiple batches according to the order of arrangement, specifically: Store all visited vertices in an array in the order of visit, and according to the preset number of threads t , and the number of vertices in the vertex stream v , determine the size of the batch b ; among them, ; Each thread determines its unique id number, , and then calculates the starting index b and id for each thread to be responsible for dividing vertices and the ending index ; When the vertices processed by the thread id exceed the end index eid or go out of the array range, the processing stops.
4. The method according to claim 1, characterized in that Step (3.2) specifically includes the following steps: (3.2.1) Calculate the penalty coefficient according to the preset hyperparameters β , the number of vertices currently partitioned into the partition, the number of partitions, and the number of vertices already partitioned , where is the number of vertices in the partition p , k is the number of partitions, is the number of vertices already partitioned, then ; (3.2.2) Take the current partition vertex v 's neighbor set N(v) and the vertex set that has been currently partitioned into p and use a high-speed intersection algorithm to find the intersection, obtaining the result set Q(p) R , and take the modulus of R as the value of the first-order attribute ; (3.2.2.1) Arrange the vertices of N(v) and Q(p) in ascending order according to the vertex numbers, and compare the magnitudes of the moduli of the two sets. Point the pointer s to the first element of the smaller set, and point the pointer to the first element of the larger set; (3.2.2.2) Set pointer s Points to element S ,pointer Points to element L ,Compare S and L Size: If S equal L , then S join in R In, and s and Both point to the next element; if S Less than L , only s Points to the next element; if S Greater than L , set the initial step size r is equal to 1, so point to position, re-compare S and L ,if S Still greater than L , then ,until S Less than L At this time, Collection Use binary search to search for a range and S Same If found, join in R , so that point to The next one; if not found, make s Point to the next element, point to location; Repeat the comparison in (3.2.2.2) until s or reaches the end of the set it belongs to, and return the set R , and proceed to step (3.2.3); (3.2.3)Traverse the set R , and count R each vertex in r the common neighbors with the current partition vertex v , and sum up all the counted numbers of common neighbors as the value of the second-order attribute ; (3.2.4) Calculate and obtain the partitioning score based on the penalty coefficient, first-order attribute and second-order attribute.
5. A social network data analysis system, characterized in that It includes: A social network graph acquisition module, which is used to store the data of the social network in the form of a social network graph, where each vertex in the social network graph represents a user. If there is a social relationship between users, there is an edge connection between the corresponding two vertices; A vertex access module, which is used to determine the vertex with the largest degree in the social network graph. According to the principle of breadth-first search with the largest degree, the neighbor vertices of this vertex are accessed in descending order of degree. Subsequently, according to the same principle, the neighbor vertices of the neighbors of this vertex are accessed until every vertex in the social network graph is accessed; A thread allocation module, which is used to arrange all the accessed vertices in the order of access, and divide all the accessed vertices into multiple batches according to the preset parallel granularity, and evenly distribute them to each thread according to the arranged order; A vertex partitioning module, which is used to partition each vertex into the corresponding partition in each thread according to the multi-order attribute perception strategy; wherein, the multi-order attribute perception strategy considers the number of neighbors of the vertex in each partition and the number of common neighbors of the vertex and the vertices in each partition, and the multi-order attribute perception strategy has a dynamic load balancing constraint on the load of each partition to ensure the load balance of multiple partitions; the vertices with similar characteristics in the social network are partitioned into the same partition to reduce the communication overhead when analyzing social network data in a distributed system; The vertex partitioning module divides each vertex into a corresponding partition in each thread according to the multi-order attribute perception strategy. Specifically: (3.1) Each thread traverses its own vertex stream. For each vertex, it then traverses all partitions; (3.2) For each vertex being partitioned currently v and each partition p calculate the partitioning score S(v,p) . The partitioning score consists of three parts: ; is the penalty coefficient. The penalty coefficient is related to the preset hyperparameter and the number of vertices already partitioned into the partition. The penalty intensity changes with the current partitioning situation to control the load balance of each partition. The value of the hyperparameter can be adjusted within a preset range to control whether to give priority to load balance or minimum cut in the partitioning; is the first-order attribute, considering the number of neighbors of vertex v in partition p ; is the second-order attribute, considering the number of common neighbors of vertex v and the vertices in partition p ; (3.3) For the vertex being partitioned currently, compare its partitioning scores in each partition, select the partition with the highest partitioning score, and assign v to the corresponding partition.
6. The system according to claim 5, wherein The vertex access module performs vertex access according to the following steps: (1.1) Find the vertex with the largest degree from the social network graph, access the vertex with the largest degree and sort its neighbors in descending order of degree; (1.2) Start accessing the neighbor vertices, and the access order is sorted by degree, and the neighbor with a higher degree is accessed first; (1.3) Mark the vertices that have been accessed during each access. When accessing a new vertex each time, check whether it has been accessed. If it has been accessed, skip this vertex; (1.4) After all the n-order neighbors of the vertex with the largest degree have been accessed, according to the access order of the n-order neighbors, perform steps (1.2) and (1.3) on each n-order neighbor, and perform the same access on the (n + 1)-order neighbors of the vertex with the largest degree until every vertex in the social network graph is accessed; where n represents the number of times steps (1.2) and (1.3) are repeatedly executed.
7. The system according to claim 5, characterized in that, The thread allocation module divides all visited vertices into multiple batches according to the order of arrangement, specifically: all visited vertices are stored in an array according to the order of access, and the number of threads is preset. t , and the number of vertices in the vertex stream v , determine the batch size b ;in, ; Each thread determines its unique id Number, , then according to b and id Calculate the starting index of each thread responsible for dividing the vertex and the end point index ; When the thread processes the vertex id Exceeding the end index eid Or if it exceeds the array range, processing stops.
8. An electronic device, characterized in that, including: At least one memory, which is used to store programs; At least one processor, which is used to execute the programs stored in the memory. When the programs stored in the memory are executed, the processor is used to execute the method according to any one of claims 1-4.
Citation Information
Patent Citations
Method and device for detecting large-scale social network communities
CN103942308A
Method for accelerating sub-graph matching based on CPU-FPGA (Central Processing Unit-Field Programmable Gate Array) hybrid platform
CN115687707A