A Cross-Social Network User Alignment Method
Through label propagation rules and node embedding representation learning, the problem of insufficient information propagation of anchor nodes is solved, and the precise matching effect of user alignment is improved.
Patent Information
- Application Number
- CN202210830120.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-15
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-07-15
AI Technical Summary
The information provided by the existing alignment method cannot be effectively propagated to further nodes, resulting in poor performance in the neighborhood precise matching when aligning users.
Design tag propagation rules to make anchor nodes continuously spread their own tag information, and learn through node embedding representation, calculate the cosine similarity across the network, and learn the representation of each user node in vector space to make up for the information loss during tag propagation.
It realizes the effective propagation and convergence of anchor node information during user alignment, and improves the precise matching effect of user alignment.
Smart Images

Figure CN115221956B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of graph neural network analysis, and in particular relates to a method for aligning users across social networks. Background Art
[0002] Network alignment is user anchor link prediction, which aims to anchor the same user in different networks, and the association between shared users across social networks is called anchor link. According to the number of shared users between networks, we can divide the network alignment task into full network alignment and partial network alignment. In the full network alignment task, the network to be learned is completely aligned and each user can be associated with an anchor link. Such networks are rare in real life. In the partially aligned network, the learned network is partially aligned and many users do not have matching objects, such as social networks. We can usually divide the network alignment task into three categories according to whether the data set is labeled: supervised learning, semi-supervised learning, and unsupervised learning.
[0003] In the supervised network alignment problem, by using the heterogeneous information in the network, a set of features can be extracted for the anchor links in the network, and the existing and non-existing anchor links can be marked as positive and negative. After training with the labeled training set, we can learn a mapping function that can be used to determine whether the anchor links in the test set are positive or negative. The algorithms belonging to supervised learning include the alignment algorithm for knowledge graph embedding, the PALE algorithm, and the DCIM algorithm. Semi-supervised learning is a model between supervised learning and unsupervised learning. It first trains a model in unlabeled samples, uses this model to process all samples, obtains a part of newly labeled samples, and then adds the newly labeled samples to the previous samples to retrain a model. The more typical semi-supervised learning models are the MAH algorithm and the IONE model. Among them, IONE proposes to maintain the similarity between followers and followers, and the structural diversity considered in IONE-D to learn network embedding representations. Unsupervised learning design models are used to process unlabeled sample sets, and there are only a large number of input objects. Since there are no prior user association pairs, only some unique representative attributes can be used to mine users with high correlation. For example, UAGA based on adversarial model, CoLink model trained by both relationships and attributes, and REGAL model using KD tree to calculate user similarity.
[0004] After the emergence of large-parameter methods such as Transformer, the attention mechanism has also been used. For example, ABNE proposed an approach based on the attention mechanism. With the development of deep learning in recent years, more and more methods have used deep learning to solve the social network alignment problem. DeepLink uses the characteristic structure of the two networks obtained through training and adopts dual learning of deep neural networks as a further alignment process. SNNA learns the projection function by generating adversarial networks, which minimizes the Wasserstein distance between the anchor distributions from two social networks. NeXtAlign believes that good negative samples should distinguish between anchor links and close node pairs that may mislead alignment, while not violating the overall alignment consistency, and constructs a new alignment scoring function that reflects multiple aspects of node embedding.
[0005] In summary, the problem with the existing technology is that the existing alignment method can only significantly improve the alignment effect of users who are close to the anchor node. When the strong information provided by the anchor node cannot be propagated to nodes farther away, it leads to poor performance in the neighborhood of precise matching during user alignment. Summary of the invention
[0006] In order to solve the above problems, the present invention proposes a method for aligning users across social networks. Figure 1 , specifically including the following steps:
[0007] S1: Obtain user nodes from two social networks. The user nodes that have been determined to be the same user in the two social networks are called anchor nodes, and the user nodes that have not been determined are called non-anchor nodes.
[0008] S2: Initialize the anchor node to a specific label and the non-anchor node to a zero label;
[0009] S3: assign a tuple to each non-anchor node, the tuple including the node's own label and a label list;
[0010] S4: Each non-anchor node starts to aggregate the labels of neighboring nodes. When the aggregation reaches the 0 label, it is ignored. When the aggregation reaches the specific label, it is added to the label list of the tuple of the non-anchor node. After the aggregation is completed, the tuple of the node forms an aggregate vector.
[0011] S5: Calculate the label similarity of the non-anchor node based on the aggregation vector, and select the label with high similarity as the label information of the non-anchor node;
[0012] S6: Repeat S4-S5 until no new labels are generated in the tuples of non-anchor nodes;
[0013] S7: Merge two different social networks, and learn the vector representation of non-anchor nodes without new labels, and update the vector representation;
[0014] S8: Calculate the similarity of the non-anchor nodes based on the updated vector representation, and align the two non-anchor nodes with the highest similarity.
[0015] Preferably, after each round of non-anchor nodes are aggregated to labels of neighbor nodes, a portion of the non-anchor nodes become new neighbor nodes, and the neighbor nodes represent nodes that have a friend relationship or a same attention relationship between the nodes.
[0016] Preferably, the similarity between non-anchor nodes is calculated based on the aggregate vector, and the calculation method includes:
[0017]
[0018] Among them, WL sim Represents the similarity between non-anchor nodes across the network; The aggregated label representation vector representing the source network; The aggregated label representation vector represents the target network; T represents the transpose operation of the vector; and ||·||2 represents the regularization operation.
[0019] Preferably, the representation learning of the vector specifically includes: randomly initializing each user node as a vector, optimizing the vector in a direction of reducing loss through a loss function, and updating the vector;
[0020] Furthermore, the loss function expression is:
[0021]
[0022] in, represents the loss function of the i-th node in the source network and the j-th node in the target network; and The vector representations of the i-node of the source network and the j-node of the target network, respectively, L ij Indicates whether the i node and the j node have the same mapping label during the aggregation process. Represents the cosine similarity function of the vector representation of the i-node of the source network and the j-node of the target network; label() represents the function of obtaining the current label of the node.
[0023] Preferably, the loss function is used to close the vector representation of non-anchor nodes with the same label information, and its expression includes:
[0024]
[0025] Among them, represents the context loss function of the i-th node in the source network and the j-th node in the target network; represents the representation vector of the i-th node in the social network after the merger of the source network and the target network; represents the representation vector of the j-th node in the social network after the merger of the source network and the target network; represents the representation vector of the input node of the j-th neighbor node in the merged network; v″ i represents the representation vector of the output node of the i-th neighbor node; C ij represents whether the j node is a neighbor node of the i node, 1 represents yes, 0 represents no; context(i) is the set of neighbor nodes of the i node; T is the transpose operation of the vector; logσ() represents the activation function based on the sigmoid function.
[0026] Preferably, calculating the similarity of non-anchor nodes according to the updated vector representation includes:
[0027]
[0028] Among them, rel(i,j) represents the node similarity between the i-th node in the source network and the j-th node in the target network; is the value of the p-th dimension of the vector representation of the i-th node in the source network; is the value of the p-th dimension of the vector representation of the j-th node in the target network; d is the dimension of the vector.
[0029] The beneficial effects of the present invention are as follows: By designing the label propagation rule, during the alignment process, the anchor node continuously broadcasts its own label information and mapping to achieve convergence, and at the same time alternates with "node embedding representation learning" to learn the representation of each user node in the vector space; and by calculating the cross-network cosine similarity, the nodes that are the most similar in both bidirectional networks are mapped to the same label, and the node embedding representation learning compensates for the information loss in the label propagation process, so that the users maintain accurate matching performance during alignment. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a schematic diagram of a cross-social network alignment method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0031] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0032] The present invention proposes a cross-social network alignment method, which specifically includes the following steps:
[0033] S1: Obtain user nodes from two social networks. The user nodes that have been determined to be the same user in the two social networks are called anchor nodes, and the user nodes that have not been determined are called non-anchor nodes.
[0034] S2: Initialize the anchor nodes with specific labels and initialize the non-anchor nodes with 0 labels.
[0035] S3: Assign a tuple to each non-anchor node. The tuple includes the label of the node itself and a label list.
[0036] S4: Each non-anchor node starts to aggregate the labels of its neighbor nodes. When aggregating to a 0 label, it is ignored. When aggregating to a specific label, it is added to the label list of the tuple of the non-anchor node. After the aggregation is completed, the tuple of the node forms an aggregation vector.
[0037] S5: Calculate the label similarity of the non-anchor nodes according to the aggregation vector, and select the label with high similarity as the label information of the non-anchor nodes.
[0038] S6: Repeat S4 - S5 until no new labels are generated in the tuples of the non-anchor nodes.
[0039] S7: Merge the two different social networks, perform vector representation learning on the non-anchor nodes where no new labels are generated, and update the vector representation.
[0040] S8: Calculate the similarity of the non-anchor nodes according to the updated vector representation, and align the two non-anchor nodes with the highest similarity.
[0041] After obtaining two different social networks, initialize the labels of each network node. The initialization rules are as follows:
[0042] First, initialize the anchor nodes with specific labels to distinguish the information of different users in the social network in the network.
[0043] Then, initialize the non-anchor nodes with label 0, which will not affect the neighboring nodes during information propagation.
[0044] After initialization, perform label propagation and mapping. That is, record the information of the neighbor nodes of each node on itself, and then re-label according to the similarity of each node. The rules of aggregation and mapping are as follows:
[0045] For each propagation step, a tuple will be assigned to each non-anchor node, which contains the label of the node itself and a multiset with the labels of the first-order neighbors. We represent the i-th user in the network as where ∑wl i = 1, and |C a | is the length of the label set, and the value at position i in the vector is set to 1. The representation vector of the node with label 0 is the zero vector.
[0046] The neighbor nodes are the nodes with a friendship relationship or the same following relationship between nodes. After each non-anchor node aggregates the labels of the neighbor nodes in each round, some non-anchor nodes become new neighbor nodes, and then repeat the aggregation and mapping process until no new neighbor nodes are generated, that is, when aggregating, no new labels are generated in the tuples of non-anchor nodes.
[0047] When performing label propagation, each node will aggregate the labels of the neighboring nodes to form a multi-hot aggregation vector, and then calculate the similarity of nodes across networks. The calculation method is as follows:
[0048]
[0049] where, WL sim represents the similarity between non-anchor nodes across networks; represents the aggregation label representation vector of the source network; represents the aggregation label representation vector of the target network; T represents the transpose operation of the vector; ||·||2 represents regularizing the aggregation label representation vectors of the source network and the target network to reduce the influence of different nodes.
[0050] Repeat the above process until the size of the label set |C a | no longer changes. This anchor-based label propagation strategy can explore the isomorphic subgraphs in the two networks and can guide subsequent graph representation learning.
[0051] The vector representation of learning nodes. Each user will get a vector representation, which is continuously learned and updated during the tagging process, and finally the representation vector of the user is obtained. The cosine similarity between users is the similarity between users. The most similar users in the cross-network are our aligned users. The specific process of learning this vector is as follows:
[0052] After each round of aggregation and hashing is completed, we perform graph embedding operations simultaneously:
[0053] We first assign three randomly initialized representation vectors to the i-th node which represent the representation vector of the node, the representation vector of its input node, and the representation vector of its output node respectively.
[0054] First of all, nodes with the same label have more similar vector representations, which is called the loss of the classifier, so that nodes with the same label are in a clique in the vector space:
[0055]
[0056]
[0057] Among them, represents the loss function of the i-th node in the source network and the j-th node in the target network; and represent the vector representations of the i-th node in the source network and the j-th node in the target network respectively, and L ij represents whether the i-th node and the j-th node have the same mapping label during the aggregation process; represents the cosine similarity function of the vector representations of the i-th node in the source network and the j-th node in the target network; label() represents the function to obtain the current label of the node.
[0058] In order to keep the nodes with similar structural backgrounds in each network close, and the anchor pairs in the network have the same representation, so as to learn a unified embedding space, we add the following loss function:
[0059]
[0060]
[0061] Among them, represents the context loss function of the i-th node in the source network and the j-th node in the target network; represents the representation vector of the i-th node in the social network after the merger of the source network and the target network; represents the representation vector of the j-th node in the social network after the merger of the source network and the target network; Represents the representation vector of the input node of the jth neighbor node in the merged network; v″ i Represents the representation vector of the output node of the i-th neighbor node; C ij Indicates whether the j node is a neighbor node of the i node, 1 represents yes, 0 represents no, that is, the relationship of first-order neighbors; context(i) is the set of neighbor nodes of the i node; T is the transpose operation of the vector; logσ() represents the activation function based on the sigmoid function.
[0062] The vector representation of the node is updated through the Adam gradient descent algorithm, including:
[0063]
[0064]
[0065] Among them, α t is the update step size; α is the initialization learning rate; are the exponential decay rates of the first-order moment and the second-order moment, respectively, and are hyperparameters set in advance; θ t is the parameter vector we want to update at time t; m t , v t They are the first-order and second-order moment estimates of the gradient obtained by differentiating the vector according to the loss function; A constant added to maintain numerical stability; following the above steps, we can update our user representation vector based on the gradient obtained by derivation of the loss function to optimize the vector to the representation we need.
[0066] This loss is achieved by bringing each user's representation vector closer to the output vector of the source node and the input vector of the target node to ensure that the embedding can retain the structural information of the network. The entire process uses the Adam gradient descent algorithm to update the vector representation of the node. When the label algorithm propagation process is completed, we use the user representation vector to calculate the similarity between users.
[0067] Calculate the similarity of non-anchor nodes based on the updated vector representation, including:
[0068]
[0069] Among them, rel(i,j) represents the node similarity between the i-th node in the source network and the j-th node in the target network; is the value of the pth dimension represented by the i-th node vector of the source network; is the value of the pth dimension represented by the vector of the jth node of the target network; d is the dimension of the vector.
[0070] The above-described embodiments have further elaborated on the object, technical solution, and advantages of the present invention. It should be understood that the above-described embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for aligning users across social networks, characterized in that, The specific steps include: S1: Obtain user nodes from two social networks. The user nodes that have been determined to be the same user in the two social networks are called anchor nodes, and the user nodes that have not been determined are called non-anchor nodes. S2: Initialize the anchor node to a specific label and the non-anchor node to a zero label; S3: assign a tuple to each non-anchor node, the tuple including the node's own label and a label list; S4: Each non-anchor node starts to aggregate the labels of neighboring nodes. When the aggregation reaches the 0 label, it is ignored. When the aggregation reaches the specific label, it is added to the label list of the tuple of the non-anchor node. After the aggregation is completed, the tuple of the node forms an aggregate vector. S5: Calculate the label similarity of the non-anchor node based on the aggregation vector, and select the label with high similarity as the label information of the non-anchor node; S6: Repeat S4-S5 until no new labels are generated in the tuples of non-anchor nodes; S7: Merge two different social networks, and learn the vector representation of non-anchor nodes without new labels, and update the vector representation; S8: Calculate the similarity of the non-anchor nodes based on the updated vector representation, and align the two non-anchor nodes with the highest similarity.
2. The method for aligning users across social networks according to claim 1, characterized in that After each round of non-anchor nodes are aggregated to the labels of neighbor nodes, label information of the friendship relationship or the same attention relationship between neighbor nodes is obtained, and they become new neighbor nodes. The neighbor nodes represent nodes with friendship relationship or the same attention relationship between nodes.
3. A method for aligning users across social networks according to claim 1, characterized in that, The label similarity of non-anchor nodes is calculated based on the aggregated vector. The calculation method includes: Among them, WL sim represents the label similarity of non-anchor nodes across networks; represents the aggregated label representation vector of the source network; represents the aggregated label representation vector of the target network; T represents the transpose operation of the vector; ||·||2 represents the regularization operation.
4. A method for aligning users across social networks according to claim 1, characterized in that, The representation learning of vectors specifically includes: randomly initializing each user node as a vector, optimizing the vector in the direction of reducing loss through the loss function, and updating the vector.
5. The cross-social network user alignment method according to claim 4, wherein The loss function expression is: Among them, represents the loss function of the $i$-th node in the source network and the $j$-th node in the target network; and respectively represent the vector representations of the $i$-th node in the source network and the $j$-th node in the target network. $L$ ij represents whether the $i$-th node and the $j$-th node have the same mapping label during the aggregation process. represents the cosine similarity function of the vector representations of the $i$-th node in the source network and the $j$-th node in the target network; $\text{label}()$ represents the function to obtain the current label of the node.
6. A method for aligning users across social networks according to claim 1, characterized in that The loss function is used to close the vector representation of non-anchor nodes with the same label information. Its expression includes: Among them, represents the context loss function of the $i$-th node in the source network and the $j$-th node in the target network; represents the representation vector of the $i$-th node in the social network after the merger of the source network and the target network; represents the representation vector of the $j$-th node in the social network after the merger of the source network and the target network; represents the representation vector of the input node of the $j$-th neighbor node in the merged network; $v″$ i represents the representation vector of the output node of the $i$-th neighbor node; $C$ ij represents whether the $j$-th node is a neighbor node of the $i$-th node, 1 represents yes, 0 represents no; $context(i)$ is the set of neighbor nodes of the $i$-th node; $T$ is the transpose operation of the vector; $logσ()$ represents the activation function based on the sigmoid function.
7. A method for aligning users across social networks according to claim 1, characterized in that Calculate the similarity of non-anchor nodes based on the updated vector representation, including: Among them, rel(i, j) represents the node similarity between the i-th node in the source network and the j-th node in the target network; is the value of the p-th dimension of the vector representation of the i-th node in the source network; is the value of the p-th dimension of the vector representation of the j-th node in the target network; d is the dimension of the vector.