A Social Network Account Alignment Method Based on Hypergraph Embedding
By using the hypergraph random walk and word embedding methods in the social network hypergraph, and the clustering based on attribute similarity calculation, the problems of insufficient user feature fusion ability and high time complexity in the existing technology are solved, and efficient cross-social network account alignment is achieved.
Patent Information
- Application Number
- CN202410520829.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-04-28
AI Technical Summary
The prior art has problems in cross-social network account alignment, which lacks the ability to fusion of multiple user features, relies on a large number of supervised matching pairs to introduce noise, and has high time complexity.
By constructing a social network hypergraph, a super-graph random walk method based on hyper-edge weight is adopted, and the word embedding method is used to embed the hypergraph nodes after clustering. After the account attribute similarity calculation is made, the alignment of cross-social network accounts is achieved.
Reduces time complexity, reduces resource requirements, improves account alignment efficiency, and can accurately match cross-social network accounts within a narrowed range.
Smart Images

Figure CN118410347B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a social network account alignment method based on hypergraph embedding. Background Art
[0002] With the rapid development of the Internet, people's lifestyles have changed greatly. In terms of social interaction, social network services have become an indispensable means of communication for contemporary people. People share their experiences, moods, statuses, opinions on certain things, etc. on social networks, which generates a large amount of valuable information. At the same time, the statements on social networks are anonymous. As long as users do not want to disclose their real information, others will not know who the user of this account is, and the same user may have corresponding accounts on multiple social networks. Therefore, it is of great significance to match a target user entity with its associated accounts and information on multiple social networks to achieve cross-social network account alignment.
[0003] In the study of social networks, traditional graph representation methods are difficult to represent user features such as non-pairwise relationships, and these features can often effectively provide information for account alignment. As a new concept proposed in recent years, a hypergraph is actually a generalization of an ordinary graph. When representing users and their relationships in a social network, an ordinary graph defaults that all relationships are binary relationships, and this representation method often oversimplifies the complexity of real data. And in the actual social network, most of the relationships between users cannot be expressed by simple binary relationships. Therefore, to address this issue, the social network hypergraph introduces the concept of hyperedges that can be connected to any number of vertices. Compared with ordinary graphs, by flexibly defining hyperedges, higher-order user features such as non-pairwise relationships in the social network can be captured. In addition, users on different social platforms may also be connected to each other through non-pairwise relationships.
[0004] The prior art endeavors to discover additional user features based on the introduction of more user data, thereby improving the accuracy of account alignment. First, extracting user features through account profiles is an intuitive technique that does not require complex algorithmic models. There is a method (Liu L, Cheung W K, Li X, et al. Aligning Users across Social Networks Using Network Embedding[C]. Ijcai, 2016:1774-1780) that uses account profiles such as the user names, genders, avatars, etc. filled in the account for account alignment. With the development of graph embedding technology, there are methods (Xu Qingting, Hong Yu, Pan Yuchen, et al. A Review of Attribute Extraction Research[J]. Journal of Software, 2023, 34(2):690-711, Narayanan A, Shmatikov V. De-anonymizing social networks[C]. 2009 30th IEEE symposium on security and privacy, 2009:173-187, Vaswani A, Shazeer N, Parmar N, et al. Attention Is All You Need[J]. arXiv e-prints, 2017:arXiv:1706.03762, Kipf T N, Welling M. Semi-supervised classification with graph convolutional networks[J]. arXiv preprint 2016:arXiv:1609.02907) that use graph embedding methods to represent different social networks respectively based on social network topology structure data, and supervised use of the topological features of users for account alignment. In addition, unique user features can also be extracted from user-generated content such as posts, comments, and forwards. There are methods (Zhang Y, Tang J, Yang Z, et al. Cosnet: Connecting heterogeneous social networks with local and global consistency[C]. Proceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining, 2015:1485-1494, Hu Sanning, Li Yuxiang. A Cross-Social Network User Matching Method Based on Multi-Source Data Integration[J]."Computer Simulation, 2021, 38(4): 352 - 355", "Man T, Shen H, Jin X, et al. Cross - domain recommendation: An embedding and mapping approach[C]. IJCAI, 2017: 2464 - 2470", "Zhan Q, Zhang J, Wang S, et al. Influence maximization across partially aligned heterogenous social networks[C]. Advances in Knowledge Discovery and Data Mining: 19th Pacific - Asia Conference, PAKDD 2015, Ho Chi Minh City, Vietnam, May 19 - 22, 2015, Proceedings, Part I 19, 2015: 58 - 69", "Zhang J, Yu P S. Pct: partial co - alignment of social networks[C]. Proceedings of the 25th international conference on World Wide Web, 2016: 749 - 759", "Zhang Z, Wen J, Sun L, et al. Efficient incremental dynamic link prediction algorithms in social network[J]. Knowledge - Based Systems, 2017, 132: 226 - 235)" use user features such as users' long - term topic interests, language styles, emoji usage habits, etc. for account alignment.
[0005] However, there are still some common defects in existing account alignment methods. First, integrating more user features in the model can help the model distinguish users. However, limited by the fusion and representation ability of the algorithm model for complex data in multi - source social networks, existing methods rarely can uniformly utilize multiple user features for account matching. Second, existing graph representation methods based on social network topology data rely on a large number of supervised matching pairs to align different social networks into the same embedding space. However, this process usually introduces unnecessary noise into the model, and supervised matching pairs are not easily obtained in the real world. And when dealing with large - scale data, the time complexity is high and a large amount of resources are consumed. Summary of the Invention
[0006] The object of the present invention is to overcome the deficiencies of the prior art and provide a method for hypergraph random walk based on hyperedge weights on the basis of constructing a social network hypergraph. After obtaining a walk sequence, the method of word embedding is used to embed the hypergraph nodes, obtain the embedding vectors of the nodes and cluster them. Subsequently, based on the calculation of the attribute similarity of the accounts within a reduced range, the alignment of cross-social network accounts is realized.
[0007] The object of the present invention is achieved by the following technical solutions: A social network account alignment method based on hypergraph embedding, comprising the following steps:
[0008] S1. Data acquisition: Select a plurality of accounts with account information on two social platforms simultaneously. Taking these accounts as the center, collect the personal information, friend relationships, and dynamic information published by the relevant social platform accounts.
[0009] S2. Construct a social network hypergraph. The social network hypergraph consists of multiple user nodes and the following four types of hyperedges:
[0010] (1) Follow relationship hyperedge: According to the collected follow relationship data between users, if there is a mutual follow relationship between any two user nodes, that is, a complete bipartite graph, it is considered that these users are on the same hyperedge.
[0011] (2) Location hyperedge: Use geographical location information to construct the hypergraph. Accounts in the same city are one hyperedge.
[0012] (3) Language hyperedge: Therefore, according to the languages used by users, accounts using the same language are one hyperedge.
[0013] (4) Interest hyperedge: Extract the themes from the user-generated content and obtain the theme feature vector α = [t1, t2,..., tk] of the user. The value ti of each dimension of α represents the probability that user A is talking about theme i; Each theme is regarded as a non-pairwise relationship, and accounts with the same interest are one hyperedge.
[0014] S3. Social network hypergraph embedding based on hypergraph random walk: On the basis of the obtained social network hypergraph, start hypergraph random walk; After obtaining the walk sequence, use the Word2Vec model for training, output the embedding vectors of each node in the hypergraph, and then use the Kmeans method to cluster the user nodes, clustering the user nodes with similar characteristics into one category.
[0015] S4. Cross - social network account alignment based on similarity calculation: Process the user descriptions, usernames, and avatar information of social network accounts, calculate the similarities between different social network platform accounts in terms of user descriptions, usernames, and avatar information respectively, and calculate the total similarity. And in each cluster obtained by clustering, only when a pair of accounts are the nodes with the highest total similarity to each other is it regarded as a successful match, so as to achieve cross - social network account alignment;
[0016] Similarity of user descriptions: Calculate the similarity s between the short text data of the user description fields; Use sentence vector technology to convert the user description text into sentence vectors, and measure the similarity of user descriptions through sentence vectors; Specifically, for user description texts A and B, use the SBERT model to first process them into token sequences [w1, w2,..., w desc and [w′1, w′2,..., w′ n , then use BERT to extract the word vectors v n of each token w i and w′ i respectively, and then use pooling technology to aggregate all word vectors v i and v′ i to obtain the final n - dimensional sentence vector representations s i and s′ i of user description texts A and B; For user description texts A and B, calculate the similarity s a between the corresponding sentence vectors through cosine similarity; a ; desc ;
[0017] Similarity of usernames: Judge their similarity by the character differences of username words, and use the Levenshtein edit distance to process the username attributes; Specifically, for usernames A and B after removing special characters, their similarity s name is calculated as:
[0018]
[0019] where operation(·) represents the minimum number of operations required to convert between two strings; max_len(·) represents the maximum length in the string;
[0020] Similarity of user avatars: Use the VGG - 16 pre - trained model to calculate the similarity of user avatars, denoted as s pic; The model extracts image features through a convolutional neural network and determines the similarity of images based on these features. For avatar image A and avatar image B, the VGG-16 model passes the image features layer by layer in its neural network and outputs their 1000-dimensional feature vectors f A and f B , and then uses cosine similarity to calculate the similarity s A between f B and f pic ;
[0021] The total similarity s attr is obtained from their average value, that is, s attr =(s name +s pic +s desc ) / 3; Based on the total similarity, the user node with the highest similarity is found within the small range after clustering and regarded as aligned.
[0022] The beneficial effects of the present invention are as follows: Based on the construction of a social network hypergraph, the present invention proposes a hypergraph random walk method based on hyperedge weights. After obtaining the walk sequence, the method of word embedding is used to embed the hypergraph nodes, obtaining the embedding vectors of the nodes and clustering them. Subsequently, within the reduced range, the alignment of cross-social network accounts is realized based on the calculation of the attribute similarity of the accounts. It can reduce the time complexity, reduce the resource requirements, and improve the alignment efficiency. Description of the Drawings
[0023] Figure 1 is the flowchart of the social network account alignment method based on hypergraph embedding of the present invention;
[0024] Figure 2 is the schematic diagram of the construction of the social network hypergraph of the present invention;
[0025] Figure 3 is the flowchart of the hypergraph random walk algorithm of the present invention;
[0026] Figure 4 is the schematic diagram of the user similarity calculation of the present invention. Detailed Embodiment
[0027] The present invention will propose a social network hypergraph embedding method based on hypergraph random walk and a cross-social network account alignment method based on attribute similarity calculation. By constructing a social network hypergraph for a large number of accounts within the social network, the large number of accounts are divided into small groups with certain similarity relationships, and within the reduced range, the alignment of cross-social network accounts is realized based on the calculation of the attribute similarity of the accounts.
[0028] The method for aligning social network accounts based on hypergraph embedding needs to solve the following key problems: (1) construction of social network hypergraphs; (2) hypergraph embedding method for social networks based on hypergraph random walk; (3) cross-social network account alignment method based on attribute similarity calculation.
[0029] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0030] As Figure 1 shown, a method for aligning social network accounts based on hypergraph embedding of the present invention includes the following steps:
[0031] S1. Data acquisition: Select multiple accounts that have account information on two social platforms simultaneously. Taking these accounts as the center, collect personal information, friendship relationships, and dynamic information published by the relevant social platform accounts.
[0032] In this embodiment, the data of two platforms, Twitter and Facebook, are taken as examples for illustration. Since the currently available datasets with built-in Twitter and Facebook account binding functions are relatively old and incomplete, it may bring errors to the experiment. Therefore, the present invention does not adopt these datasets. And at the same time, since there are some platform accounts on LinkedIn that have filled in both Facebook account information and Twitter account information. Therefore, we can consider that the corresponding relationship of this part of Twitter and Facebook accounts is accurate. Then the present invention sets the selection target as LinkedIn accounts in the European region, and a total of 200 accounts that have both Twitter and Facebook account information are selected. Taking these accounts as the center, collect personal information, friendship relationships, and tweet and dynamic release information of the relevant Twitter accounts and Facebook accounts. A total of 1794 Twitter accounts' personal information and their friendship relationships, as well as 1,673,955 tweets they have published in total, are crawled; 1120 Facebook accounts' personal information and their friendship relationships, as well as 85,911 Facebook dynamics they have published, are crawled.
[0033] S2. Construct a social network hypergraph, which consists of multiple user nodes and the following four types of hyperedges:
[0034] (1) Follow relationship hyperedge: According to the collected follow relationship data between users, if there is a mutual follow relationship between any two user nodes, that is, a complete bipartite graph, it is considered that these users are on the same hyperedge;
[0035] (2) Location hyperedge: Users will publicly disclose their geographical location information on their account home pages. The geographical location attributes of the accounts of the same person are usually consistent. Therefore, geographical location information can be used to construct a hypergraph, and the accounts in the same city form a hyperedge;
[0036] (3) Language hyperedge: Users will fill in the languages they use in their account profiles. The languages used by the accounts of the same person should be consistent. Therefore, according to the languages used by users, the accounts using the same language form a hyperedge;
[0037] (4) Interest hyperedge: Extract the themes from the user-generated content and obtain the user's theme feature vector α = [t1, t2, …, tk]. The value ti of each dimension of α represents the probability that user A is talking about theme i; regard each theme as a non-pairwise relationship, and the accounts with the same interest form a hyperedge;
[0038] S3. Social network hypergraph embedding based on hypergraph random walk: Based on the obtained social network hypergraph, start hypergraph random walk; after obtaining the walk sequence, use the Word2Vec model for training, output the embedding vectors of each node in the hypergraph, and then use the Kmeans method to cluster the user nodes, clustering the user nodes with similar features into one category;
[0039] The hypergraph random walk includes the following sub-steps:
[0040] S31. Traverse the user nodes in the hypergraph, select a user node as the starting node, and perform node walk; and initialize the walk sequence;
[0041] S32. Judge whether the number of node walks is less than the preset threshold n. If so, execute step S32; otherwise, output the current walk sequence, then return to S31, select the next user node as the starting node and perform node walk;
[0042] S33. Judge whether the length of the walk sequence is less than the preset value m. If so, execute step S34; otherwise, increment the number of node walks by 1 and return to step S32;
[0043] S34. Calculate the transition probability of the node hyperedge, and randomly select the next hyperedge according to the transition probability: When calculating the transition probability of the hyperedge, it is considered that different hyperedges do not contribute equally to the user similarity measurement. Therefore, weights need to be assigned to the hyperedges according to prior knowledge, where the weight of the attention relationship hyperedge > the weight of the location hyperedge = the weight of the language hyperedge > the weight of the interest hyperedge; for the hyperedge e where the current node is located i the transition probability p i is calculated as shown in the following formula:
[0044]
[0045] Among them, w i is the weight of the hyperedge e i , n is the number of hyperedges where the node is located, and w k is the weight corresponding to each type of hyperedge;
[0046] S35. Calculate the transition probability of the nodes within the hyperedge, and randomly select the next node according to the transition probability: When calculating the transition probability of the user nodes within the hyperedge, based on the co-occurrence times of the current node and other nodes within the hyperedge in other hyperedges, and comprehensively calculate the weights of each candidate node for the next hop considering the types of co-occurring hyperedges. The specific calculation of the node weight is shown in the following formula:
[0047]
[0048] Among them, n represents the number of hyperedges where the node u i co-occurs with the current node, and w(e j ) represents the weight assigned to the corresponding hyperedge e j . After calculating the weights of the candidate user nodes respectively, determine the transition probability according to the magnitude of the weights. For the node u i , the transition probability p i is calculated as shown in the following formula:
[0049]
[0050] Among them, w i is the weight of the node u i , n is the number of nodes included in the current hyperedge, and w k is the weight corresponding to each node;
[0051] S36. Add the selected node to the random walk sequence, and return to step S34;
[0052] The specific implementation of the hypergraph random walk is as Figure 3 shown, where the number of node random walk times n and the length of the random walk sequence m can be determined according to the data situation.
[0053] S4. In S3, the user node embedding of the social network hypergraph and the clustering of similar users are realized, narrowing the scope of account alignment search. The present invention further proposes a cross-social network account alignment method based on similarity calculation: Process the user descriptions, user names, and avatar information of social network accounts, calculate the similarities between the accounts on different social network platforms in terms of user descriptions, user names, and avatar information respectively, and calculate the total similarity. And in each cluster obtained by clustering, only when a pair of accounts are each other's nodes with the highest total similarity is it considered a successful match, so as to achieve the alignment of cross-social network accounts; The specific process is as Figure 4 shown.
[0054] Similarity of user descriptions: Both Twitter and Facebook provide user description fields in the account for users to fill in. According to the platform regulations, it is usually short text data. To calculate the similarity s between the short text data in the user description fields desc , the sentence vector technology is used to convert the user description text into sentence vectors, and the similarity of user descriptions is measured by the sentence vectors; specifically, for the user description texts A and B, the SBERT model is used to first process them into token sequences [w1, w2,..., w n and [w′1, w′2,..., w′ n , then use BERT to extract the word vectors v i of each token w i and w′ i and v′ i , and then aggregate all the word vectors v i and v′ i through the pooling technology to obtain the final n-dimensional sentence vector representations s a and s′ a of the user description texts A and B; for the user description texts A and B, the cosine similarity is used to calculate the similarity s desc between the corresponding sentence vectors;
[0055] Similarity of usernames: The usernames in the account are usually proper nouns such as names and nicknames, and it is difficult for existing embedding models to embed and represent them. However, precisely because proper nouns such as names and nicknames have good uniqueness, the present invention directly judges their similarity through the differences in characters of the username words, and uses the Levenshtein edit distance to process the username attribute; specifically, for the usernames A and B after removing special characters, their similarity s name is calculated as:
[0056]
[0057] where operation(·) represents the minimum number of operations required to convert between two strings; max_len(·) represents the maximum length in the string; since the number of operations required for conversion is always less than or equal to the maximum length in the string, the similarity calculated by the Levenshtein edit distance is always within [0, 1].
[0058] Similarity of user avatars: The VGG-16 pre-trained model is used to calculate the similarity of user avatars, denoted as s pic; The model extracts image features through a convolutional neural network and judges the similarity of images based on these features. For avatar image A and avatar image B, the VGG-16 model transmits the image features layer by layer in its neural network and outputs their 1000-dimensional feature vectors f A and f B , and then uses the cosine similarity to calculate the similarity s A between f B and f pic ;
[0059] The present invention believes that self-introduction, username, and avatar are equally important for judging the similarity of accounts. Therefore, the total similarity s attr is obtained from their average value, that is, s attr =(s name +s pic +s desc ) / 3; The user node with the highest similarity is searched for within a small range after clustering based on the total similarity and regarded as alignment.
[0060] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention according to the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A social network account alignment method based on hypergraph embedding, characterized in that: The following steps are involved: S1. Data acquisition: Select multiple accounts that have account information on both social platforms, and collect personal information, friend relationships, and dynamic information posted by the relevant social platform accounts with these accounts as the center; S2. Build a social network hypergraph, which consists of multiple user nodes and the following four types of hyperedges: (1) Attention relationship hyperedge: Based on the collected attention relationship data between users, it is defined that if there is a mutual attention relationship between any two user nodes, that is, a bidirectional complete graph, these users are considered to be on the same hyperedge; (2) Location hyperedge: Use geographic location information to construct a hypergraph, where accounts in the same city are a hyperedge; (3) Language hyperedge: Therefore, according to the language used by the user, accounts using the same language are a hyperedge; (4) Interest hyperedge: Topic extraction is performed on user-generated content, and the user's topic feature vector α = [t1, t2, ..., tk] is obtained. The value ti of each dimension of α represents the probability that user A is talking about topic i. Each topic is regarded as a non-paired relationship, and accounts with the same interest are a hyperedge. S3. Social network hypergraph embedding based on hypergraph random walk: Based on the obtained social network hypergraph, start hypergraph random walk; after obtaining the walk sequence, use the Word2Vec model for training, output the embedding vector of each node in the hypergraph, and then use the Kmeans method to cluster the user nodes, and cluster the user nodes with similar characteristics into one category; S4. Alignment of cross-social network accounts based on similarity calculation: Process the user description, username and avatar information of social network accounts, calculate the similarity between the user description, username and avatar information of accounts on different social network platforms, and calculate the total similarity; and in each cluster obtained by clustering, only when a pair of accounts are the nodes with the highest total similarity to each other, is it considered a successful match, so as to achieve alignment of cross-social network accounts; Similarity of user description: Calculate the similarity between the short text data of the user description field s desc ; Use sentence embedding technology to convert user description text into sentence embedding, and use sentence embedding to measure the similarity of user descriptions; Specifically, for user description texts A and B, use the SBERT model to first process them into word sequences [w1,w2,...,w n ] and [w′1,w′2,...,w′ n ], and then use BERT to extract each word w i and w′ i The word vector v i and v′ i , and then aggregate all word vectors v through pooling technology i and v′ i , thus obtaining the final n-dimensional sentence vector representation s of user description texts A and B a and s′ a ; For user description texts A and B, the similarity s between the corresponding sentence vectors is calculated by cosine similarity desc ; Similarity of usernames: The similarity of usernames is determined by the differences in characters between the words in the usernames. The Levenshtein edit distance is used to process the username attributes. Specifically, for usernames A and B with special characters removed, the similarity between them is s name Calculated as: Where operation(·) represents the minimum number of operations required to convert between two strings; max_len(·) represents the maximum length of a string; User avatar similarity: The VGG-16 pre-trained model is used to calculate the similarity of user avatars, denoted as s pic ; The model extracts image features through convolutional neural networks and judges the similarity of images based on these features; For avatar images A and B, the VGG-16 model passes image features layer by layer in its neural network and outputs their 1000-dimensional feature vector f in the last layer A and f B , and then use cosine similarity to calculate f A and f B The similarity between pic ; The total similarity s attr The average value of them is s attr =(s name +s pic +s desc ) / 3; based on the total similarity, the user node with the highest similarity is found in the small range after clustering, which is considered as alignment.
2. The method for aligning social network accounts based on hypergraph embedding according to claim 1, characterized in that: The hypergraph random walk includes the following sub-steps: S31, traverse the user nodes in the hypergraph, select a user node as the starting node, perform node walk, and initialize the walk sequence; S32, determining whether the number of node wandering times is less than a preset threshold value n, if yes, executing step S32; Otherwise, the current walking sequence is output, and then the process returns to S31, and the next user node is selected as the starting node for node walking; S33, determine whether the length of the walking sequence is less than the preset value m, if so, execute step S34, otherwise, increase the number of node walking times by 1 and return to step S32; S34. Calculate the transition probability of the node hyperedge, and randomly select the next hyperedge according to the transition probability: When calculating the transition probability of the hyperedge, it is considered that the contribution of different hyperedges to the user similarity measurement is not completely consistent. Therefore, it is necessary to assign weights to the hyperedges based on prior knowledge, focusing on the hyperedge weight of the relationship > the hyperedge weight of the position = the hyperedge weight of the language > the hyperedge weight of the interest; for the hyperedge e where the current node is located i The transition probability p i The calculation is as follows: Among them, w i is the hyperedge e i The weight of the node, n is the number of hyperedges where the node is located, and w k is the weight corresponding to each hyperedge; S35. Calculate the transfer probability of the node in the hyperedge, and randomly select the next node according to the transfer probability: When calculating the transfer probability of the user node in the hyperedge, the weight of each candidate node of the next hop is calculated according to the number of co-occurrences of the current node and other nodes in the hyperedge in other hyperedges, and the types of co-occurring hyperedges. The specific calculation of the node weight is shown in the following formula: Where n represents the node u i The number of hyperedges that co-occur with the current node, w(e j ) is represented by the corresponding hyperedge e j After calculating the weights of candidate user nodes respectively, the transfer probability is determined according to the weights. i Transition probability p i The calculation is as follows: Among them, w i For node u i The weight of n is the number of nodes contained in the current hyperedge, and w k The weight corresponding to each node; S36. Add the selected node to the walking sequence and return to step S34.
Citation Information
Patent Citations
Method for linking accounts in OSNs (On-line Social Networks)
CN105741175A
Long text matching method based on graph convolution
CN116304749A