Attribute social network alignment method based on attribute and structure consensus learning

By adopting a consensus-based learning method in attribute social networks, learning the attributes and structural characteristics of users and mitigating deviations through consensus constraints, the problem of user feature representation similarity relationship damage caused by attribute and structural consistency deviation in the prior art is solved, and a more accurate and consistent network alignment effect is achieved.

CN120013698APending Publication Date: 2025-05-16Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510040247.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing attribute social network alignment method fails to effectively consider the consistency bias between attributes and structures, resulting in the destruction of the similarity relationship of user feature representations in different social platforms.

Method used

Using a method based on attribute and structural consensus learning, the user's attributes and structural characteristics are learned through graph autoencoder, reconstruction loss constraints and consistency constraints are established, and the deviations between attributes and structures are alleviated through consensus constraints.

Benefits of technology

Effectively inferring correspondence between users across networks improves the accuracy and consistency of network alignment, and is better than a variety of benchmarking methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013698A_ABST
    Figure CN120013698A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of social networks, and particularly relates to an attribute social network alignment method based on attribute and structure consensus learning. The method comprises the following steps: firstly, learning double-view representation of user attributes and structures by using an auto-encoder and a graph auto-encoder, and establishing a reconstruction loss constraint based on an encoding-decoding framework; secondly, based on a consistency hypothesis, that is, the same user always shows similar hobbies and interests or local neighborhood structures in different social networks, double consistency constraints based on attributes and structures are established respectively; meanwhile, in order to relieve the influence caused by deviation between attributes and structures, consensus constraints based on common information mining are provided, and mutual unification of attribute consistency and structure consistency is enhanced. A large number of experiments on a three-attribute social network data set show that the method provided by the invention is superior to various advanced reference methods. Meanwhile, an ablation experiment further explains the rationality and effectiveness of the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of social networks, and in particular relates to an attribute social network alignment method based on attribute and structure consensus learning. Background Art

[0002] Online social networks are becoming more and more popular and gradually penetrate into every aspect of people's lives. Different social networks provide different types of services, and people often participate in multiple different social networks at the same time according to their personalized needs. For example, users follow social hot spots and news trends in the form of tweets on Twitter, share life dynamics through pictures, geographic locations, etc. on Foursquare, and also post photos and posts on Instagram. Social network alignment aims to discover the correspondence between different identities of the same user in multiple social networks, also known as anchor link prediction, cross-social network user identity recognition, etc., which is of great significance for tasks such as cross-platform friend recommendation and cross-network information dissemination.

[0003] As an important part of social network analysis, social network alignment has attracted extensive attention from scholars in recent years. Social network alignment methods can be roughly divided into three categories. The first type of method is based on user attribute information, such as user name, avatar, location, etc. This type of method extracts user feature representation from self-reported information, and its performance mainly depends on the accuracy of attribute information. The second type of method is based on user behavior information, and obtains user feature representation from user behavior data, such as tweets, trajectories, user writing style, etc. This type of method is mainly limited by whether the user behavior information is rich. The third type of method is based on social relationship structure, and extracts structural features from user social relationships to characterize users. These methods based on single data attributes have solved the problem of social network alignment to a certain extent, but cannot fully utilize the multi-dimensional information in social networks to characterize user characteristics. In recent years, many scholars have tried to combine user attribute information or user behavior information with social relationship structure information, and have carried out related research based on the combination of multi-dimensional information.

[0004] Existing attribute social network alignment methods usually simply concatenate attributes and structures, or use attribute information as the initial features of nodes to achieve the fusion of attributes and structures through graph convolutional neural networks. Zhang et al. (Zhang S, Tong H, Xia Y, et al. Nettrans: Neural cross-network transformation [C] / / Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 2020: 986-996.) proposed the NetTrans method, which transfers and aggregates attribute information on the graph structure based on a graph convolutional neural network to achieve the characterization of user characteristics. Yan et al. (Yan Y, Zhang S, Tong H. Bright: A bridging algorithm for network alignment [C] / / Proceedings of the web conference. 2021: 3907-3917.) proposed the Bright method, which first takes the one-hot encoding of known anchor links as the basis, constructs a unified structural vector representation space by restarting random walks, and processes attribute information in a similar way to NetTrans. Finally, it concatenates attribute features and structural features as the final representation of the user.

[0005] Existing methods use different ways to fuse user attributes and structural characteristics, and have achieved good network alignment results. However, due to the diverse and personalized services and functions of different social platforms, there are certain deviations in the attributes and structures of the same user on different social platforms, that is, users who are similar in attribute characteristics are not necessarily similar in structural characteristics. For example, if users who are originally similar in attributes have dissimilar local neighborhood structures, then the aggregation of graph convolutional neural networks will destroy the original similarity relationship. Similarly, users who are originally dissimilar in attributes will lead to the same phenomenon even if they have similar local neighborhood structures. Existing methods directly splice or use graph neural networks to fuse attributes and structures, ignoring the impact of the consistency deviation between the two. Summary of the invention

[0006] In order to solve the problem that the existing attribute social network alignment methods have not considered the consistency deviation between attributes and structures, the present invention provides an attribute social network alignment method based on attribute and structure consensus learning.

[0007] The present invention specifically adopts the following technical solutions:

[0008] The present invention provides an attribute social network alignment method based on attribute and structure consensus learning ( C onsensus L earning between A Attribute and S This method first uses an autoencoder to learn a dual-view representation of user attributes and structure, and establishes a reconstruction loss constraint based on the encoding-decoding framework. Then, based on the consistency assumption that the same user often shows similar interests and local neighborhood structures on different social platforms, dual consistency constraints based on attributes and structures are established respectively. At the same time, in order to reduce the impact of the deviation between attributes and structures, a consensus constraint based on common information mining is proposed to unify attribute consistency and structural consistency.

[0009] The attribute social network alignment method based on attribute and structure consensus learning of the present invention comprises the following steps:

[0010] Step 1: Based on the autoencoder, the attribute feature representation of the attribute social network user is learned and the attribute reconstruction loss is established. The autoencoder is an autoencoder based on a multi-layer perceptron model.

[0011] Step 2: Extract the structural features of attribute social network users based on the graph autoencoder and establish a structural reconstruction loss. The graph autoencoder uses a graph convolutional neural network as an encoder. The structural reconstruction loss is established by minimizing the difference between the original adjacency matrix and the reconstructed adjacency matrix.

[0012] Step 3: Establish structural consistency constraints or attribute consistency constraints. Take the known anchor links as natural positive samples, construct negative samples in a random way, and establish structural consistency constraints and attribute consistency constraints based on the distance relationship between positive and negative samples.

[0013] Step 4: Establish consensus constraints between attribute feature representation and structural feature representation. In order to alleviate the impact of consistency deviation between attributes and structures on network alignment, consensus constraints are established by mining the common information between the two to enhance the mutual unification of attribute consistency and structural consistency.

[0014] Step 5: Parameter update and method optimization. The optimization goal of the method is established based on the attribute reconstruction loss constraint, structure reconstruction loss constraint, attribute consistency constraint, structure consistency constraint and consensus constraint, and convergence is achieved through iterative training.

[0015] Step 6: Network alignment. After the method converges, the attribute feature representation and the structural feature representation are concatenated as the final representation of the user. Based on the similarity measurement, the network alignment problem is transformed into a ranking retrieval problem.

[0016] As an implementation method, step 1 includes:

[0017] In order to extract dense and continuous attribute feature representation from the original attribute information, a multilayer perceptron (MLP) model is first used as an encoder to extract the user attribute feature representation Z U :

[0018] Z U =encoder U (X U ) (1)

[0019] Among them, Z is used to represent the learned user feature representation, X represents the attribute feature matrix, U represents the attribute feature, and X U Represents the initial attribute feature representation. Similarly, another MLP is used as a decoder to reconstruct the original attributes:

[0020]

[0021] in, Represents the reconstructed attribute feature representation.

[0022] Then, based on the mean square error loss function, the attribute reconstruction loss is established by minimizing the difference between the reconstructed attribute and the original attribute:

[0023]

[0024] Among them, ||·|| F represents the Frobenius norm, represents the attribute reconstruction loss.

[0025] As an implementation method, step 2 includes:

[0026] Step 2.1: Structural feature encoding. Use the graph convolutional neural network GCN as the local structural feature encoder, with the adjacency matrix (A s ) and the adjacency matrix of the target network (A t ) as input, and the transmission and aggregation of user information is realized through convolution operation:

[0027]

[0028] Among them, A represents the adjacency matrix of the attribute social network, s and t represent the source network and the target network, and H (l-1) and H (l)Respectively represent the input and output of the lth layer GCN, l represents the current GCN layer, Indicates the addition of a self-connected adjacency matrix, I N is the unit matrix, N represents the number of users in the social network, express The degree matrix of represents the weight matrix of the lth layer. σ(·) is the activation function, such as ReLU. Therefore, we can get the user structure feature representation Z V :

[0029] Z V =encoder V (X V ,A) (5)

[0030] Among them, V represents the structural characteristics, X V Represents the initial structural feature representation, obtained by random initialization.

[0031] Step 2.2: Structural reconstruction loss. Use the inner product decoder to reconstruct the relationship structure between the source network and the target network, and establish the structural reconstruction loss by minimizing the difference between the reconstructed adjacency matrix and the original adjacency matrix. The reconstructed adjacency matrix is ​​expressed as follows:

[0032]

[0033] in, To reconstruct the adjacency matrix, Represents Z V The transpose of , σ(·) is the activation function. The structure reconstruction loss based on the cross entropy loss function is:

[0034]

[0035] in, represents the structural reconstruction loss, y and Represent the adjacency matrix A and the reconstructed adjacency matrix respectively. The elements in , N represents the number of users in the social network.

[0036] The reconstruction loss constraints corresponding to step 1 and step 2 are as follows:

[0037]

[0038] in, Represents the sum of the two social network attribute reconstruction losses and the structure reconstruction loss.

[0039] As an implementation method, step 3 includes:

[0040] Step 3.1: Construct positive and negative samples. Link the anchors with known correspondences As a natural positive sample, the anchor node v is randomly i s Find K users in the target network Thus construct K negative samples Similarly, it can be an anchor node Find K users in the source network Thus, K negative samples are constructed again The set of positive samples and negative samples is called the training sample set, denoted by

[0041] in, Represents the set of users in a social network. i, j, m, n are used to distinguish different users. represents the set of known anchor links between two networks, Represents the set of positive and negative samples used for model training.

[0042] Step 3.2: Establish contrast loss. For any sample in the training sample set The contrast loss is expressed as follows:

[0043]

[0044] in, represents the contrast loss, represents the set of positive and negative samples used for model training, and y′ represents the sample The label of the positive sample is 1, the negative sample is 0, and m and n are used to represent two different users; Representation sample The Euclidean distance of Indicates the number of samples in the training sample set; margin is the set threshold.

[0045] The corresponding constraints for step 3 are as follows:

[0046]

[0047] in, Represents a consistency constraint.

[0048] As an implementation method, step 4 includes:

[0049] First, L2 regularization is used to regularize the attribute feature representation and the structural feature representation, and then the similarity matrix between users under different feature views can be obtained:

[0050]

[0051] Among them, S Urepresents the attribute feature similarity matrix, S V represents the structural feature similarity matrix, Z Unor and Z Vnor Z U and Z V L2 regularization.

[0052] Then, based on the attribute feature similarity matrix S U and the structural feature similarity matrix S V It should be a similar assumption. A consensus constraint is established based on the mean square error function. The consensus constraint is expressed as follows:

[0053]

[0054] in, Represents a consensus constraint.

[0055] As an implementation method, step 5 includes:

[0056] Combining formulas (8), (10) and (12), the overall optimization objective of the method of the present invention is established, and the Adam optimizer is used to minimize the overall loss function

[0057]

[0058] in, represents the overall loss function.

[0059] The error back-propagation method is applied to realize the iterative optimization of the proposed method.

[0060] As an implementation method, step 6 includes:

[0061] After the training of the method converges, the attribute feature representation and the structural feature representation are concatenated as the final feature representation of the user:

[0062]

[0063] Based on similarity measurement, the network alignment problem is transformed into a ranking retrieval problem. For users in the source (target) network Based on the Euclidean distance, the most similar k users are found in the target (source) network as its top-k candidate set.

[0064] The beneficial effects of the present invention are:

[0065] The present invention learns dual-view representations of user attributes and structures respectively, and infers the correspondence between users across networks based on dual consistency constraints of attributes and structures.

[0066] The present invention points out the impact of the consistency deviation between attributes and structures on the social network alignment problem, and proposes a consensus constraint based on common information mining to promote the mutual unification of attribute consistency and structural consistency.

[0067] Experimental results on three attribute social network datasets show that the proposed method outperforms a variety of state-of-the-art baseline methods. At the same time, the results of ablation experiments illustrate the rationality and effectiveness of the proposed method. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 Schematic diagram of the framework of the attributed social network alignment system based on attribute and structure consensus learning.

[0069] Figure 2 Comparison of experimental results on three attribute social network datasets; (a) the change of precision@k with k on Acm-Dblp, (b) the change of precision@k with k on Online-Offline, (c) the change of precision@k with k on Cora1-Cora2, (d) the change of MAP indicators on the three datasets.

[0070] Figure 3 Experimental results with the change of anchor link ratio; (a) precision@1 with the change of anchor link ratio, (b) MAP with the change of anchor link ratio.

[0071] Figure 4 The experimental results show how the number of negative samples changes; (a) the precision@1 changes with the number of negative samples, and (b) the MAP changes with the number of negative samples. DETAILED DESCRIPTION

[0072] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0073] Different deep graph representation learning models mainly include methods based on autoencoders and methods based on graph neural networks. The autoencoder consists of two parts: an encoder and a decoder. The encoder maps the input data to a low-dimensional code in the latent space, and the decoder remaps the code to an output similar to the original input, and learns an efficient representation of the input data by minimizing the reconstruction error. The Graph Auto-encoder (GAE) uses the Graph Convolution Neural Networks (GCN) as an encoder to obtain the user's feature representation from the graph data information by reconstructing the graph structure. The present invention uses an autoencoder based on a multi-layer perceptron model to learn user attribute features and uses a graph autoencoder to capture user structural features.

[0074] The present invention uses represents a property social network, where represents the set of N users in a social network, represents the set of connections between users in a social network, and X represents the attribute feature matrix. Representing attributes of social networks The adjacency matrix of user v i and user v j is connected, then a ij =1; otherwise a ij = 0. Given two social networks, denoted as source networks and target network use Represents the set of known anchor links between two networks.

[0075] If not otherwise specified, in the present invention, lowercase letters s and t in superscripts are used to indicate the source network and the target network, uppercase letters U and V in subscripts represent attribute features and structural features, respectively, and uppercase letter T in superscripts represents the transpose of a matrix (vector).

[0076] Example 1

[0077] like Figure 1 As shown, the attribute social network alignment method based on attribute and structure consensus learning in this embodiment includes three main modules:

[0078] (1) Structural feature learning module. This module uses graph convolution as a structural feature encoder to map the source network and the target network into a structural vector representation space to obtain the user structural feature representation.

[0079] (2) Attribute feature learning module. This module uses a multi-layer perceptron model (MLP) as an attribute feature encoder to extract user attribute feature representation from sparse original user attribute information. The user attribute information includes user name, avatar, location, etc.

[0080] (3) Optimization target establishment module. This module contains reconstruction loss constraints, consistency constraints, and consensus constraints. The purpose of reconstruction loss is to ensure that the learned structure or attribute feature representation does not deviate from the original structure or attribute characteristics; the consistency constraint based on contrast loss aims to shorten the distance between positive samples and increase the distance between negative samples, thereby enhancing the distinguishability of the structure or attribute feature representation; the consensus constraint based on mean square error loss alleviates the impact of consistency deviation between attributes and structures by mining the common information between the two.

[0081] The above-mentioned attribute social network alignment method based on attribute and structure consensus learning includes the following steps:

[0082] Step 1: Attribute feature learning based on autoencoder. Figure 1 As shown in ①, an autoencoder based on a multi-layer perceptron model is used to learn attribute feature representation and establish attribute reconstruction loss.

[0083] Step 2: Structural feature learning based on graph autoencoder. Figure 1 As shown in ②, a graph convolutional neural network is used as an encoder to extract user structural features, and a structural reconstruction loss is established by minimizing the difference between the original adjacency matrix and the reconstructed adjacency matrix.

[0084] Step 3: Establish consistency constraints. Take the known anchor links as natural positive samples, construct a certain proportion of negative samples in a random way, and establish structural and attribute consistency constraints based on the distance relationship between positive and negative samples, such as Figure 1 ③As shown.

[0085] Step 4: Establish consensus constraints. In order to alleviate the impact of consistency deviation between attributes and structures on network alignment, consensus constraints are established by mining the common information between the two to enhance the mutual unity of attribute consistency and structural consistency, such as Figure 1 ④As shown.

[0086] Step 5: Parameter update and method optimization. The optimization goal of the method is established based on the attribute reconstruction loss constraint, structure reconstruction loss constraint, attribute consistency constraint, structure consistency constraint and consensus constraint, and convergence is achieved through iterative training, such as Figure 1 ⑤As shown.

[0087] Step 6: Network alignment. After the method converges, the attribute feature representation and the structural feature representation are concatenated as the final representation of the user. Based on the similarity metric, the network alignment problem is transformed into a ranking retrieval problem, such as Figure 1 ⑥As shown.

[0088] Next, the implementation process of each step of the attribute social network alignment method based on attribute and structural consensus learning is elaborated in detail.

[0089] For step 1: attribute feature learning based on autoencoder, it includes:

[0090] The original attribute information of users in social networks is very sparse and contains a lot of noisy noise information. In order to extract dense and continuous attribute feature representation from the original attribute information, a multilayer perceptron model (MLP) is first used as an encoder to extract the user attribute feature representation Z. U :

[0091] Z U =encoder U (X U ) (1)

[0092] Among them, X U Represents the initial attribute feature representation.

[0093] Similarly, use another MLP as a decoder to reconstruct the original attributes:

[0094]

[0095] in, Represents the reconstructed attribute feature representation.

[0096] Then, based on the mean square error loss function, the attribute reconstruction loss is established by minimizing the difference between the reconstructed attribute and the original attribute:

[0097]

[0098] Among them, ||·|| F represents the Frobenius norm, represents the attribute reconstruction loss.

[0099] For step 2: structural feature learning based on graph autoencoder, it includes:

[0100] 2.1 Structural feature encoding. Use graph convolutional neural network GCN as the local structural feature encoder, with the adjacency matrix A of the source network and the target network s and A t As input, the user information is transferred and aggregated through convolution operation:

[0101]

[0102] Among them, H (l-1)and H (l) Respectively represent the input and output of the l-th layer GCN, Indicates the addition of a self-connected adjacency matrix, I N is the unit matrix, express The degree matrix of represents the weight matrix of the lth layer. σ(·) is the activation function, such as ReLU. Therefore, we can get the user structure feature representation Z V :

[0103] Z V =encoder V (X V ,A) (5)

[0104] Among them, X V Represents the initial structural feature representation, obtained by random initialization.

[0105] 2.2 Structural reconstruction loss. The inner product decoder is used to reconstruct the relationship structure between the source network and the target network, and the structural reconstruction loss is established by minimizing the difference between the reconstructed adjacency matrix and the original adjacency matrix. The reconstructed adjacency matrix is ​​expressed as follows:

[0106]

[0107] in To reconstruct the adjacency matrix, Represents Z V The transpose of , σ(·) is the activation function. The structure reconstruction loss based on the cross entropy loss function is:

[0108]

[0109] in, represents the structural reconstruction loss, y and Represent the adjacency matrix A and the reconstructed adjacency matrix respectively. The elements in , N represents the number of users in the social network.

[0110] Since both attribute feature representation learning and structural feature representation learning are based on the encoding-decoding architecture, and for a clearer description, the reconstruction loss constraints corresponding to steps 1 and 2 are summarized as follows:

[0111]

[0112] in, Represents the sum of the two social network attribute reconstruction losses and the structure reconstruction loss.

[0113] For step 3: establish consistency constraints, it includes:

[0114] 3.1 Construct positive and negative samples. Link anchors with known corresponding relationships As a natural positive sample, the anchor node is randomly Find K users in the target network Thus construct K negative samples Similarly, it can be an anchor node Find K users in the source network Thus, K negative samples are constructed again The set of positive samples and negative samples is called the training sample set, denoted by

[0115] 3.2 Establish contrast loss. Figure 1 As shown in ③, for any sample in the training sample set The contrast loss is expressed as follows:

[0116]

[0117] in, Represents the set of positive and negative samples used for model training, i, j, m, n are used to distinguish different users, and y′ represents the sample The label of , positive sample is 1, negative sample is 0; Representation sample The Euclidean distance of Indicates the number of samples in the training sample set; margin is the set threshold.

[0118] The method of the present invention establishes attribute consistency constraints from two perspectives: attribute and structure. and structural consistency constraints like Figure 1 As shown by the red double-headed arrow in ③. The corresponding constraints of step 3 are summarized as follows:

[0119]

[0120] in, Represents consistency constraints;

[0121] For step 4: establish consensus constraints, including:

[0122] In order to mitigate the impact of the deviation between user attributes and structure, a consensus constraint based on common information mining is designed to enhance the mutual unity of attribute consistency and structural consistency. In order to mine the common information between attributes and structures, the similarity of paired users in the two feature views should be similar. First, L2 regularization is used to regularize the attribute feature representation and the structural feature representation, and then the similarity matrix between users under different feature views can be obtained:

[0123]

[0124] Among them, Z Unor and Z Vnor Z U and Z V L2 regularization.

[0125] Then, based on the attribute feature similarity matrix S U and the structural feature similarity matrix S V Similar assumptions are used to establish consensus constraints based on the mean square error function, such as Figure 1 As shown by the blue bidirectional arrows in ④, consensus constraints It is expressed as follows:

[0126]

[0127] For step 5: parameter update and method optimization, including:

[0128] Combining formulas (8), (10) and (12), the overall optimization objective of the method of the present invention is established, and the Adam optimizer is used to minimize the overall loss function

[0129]

[0130] The error back propagation method is applied to realize the iterative optimization of the proposed method, such as Figure 1 ⑤As shown.

[0131] For Step 6: Network Alignment, include:

[0132] After the training of the method converges, the attribute feature representation and the structural feature representation are concatenated as the final feature representation of the user:

[0133]

[0134] Based on similarity measurement, the network alignment problem is transformed into a ranking retrieval problem, such as Figure 1 ⑥ As shown in Figure 6. For users in the source (destination) network Based on the Euclidean distance, the most similar k users are found in the target (source) network as its top-k candidate set.

[0135] Example 2 Application Example

[0136] In order to verify the effectiveness of the proposed method, experiments were carried out on three commonly used attribute social network datasets, and a detailed comparative analysis was performed with multiple benchmark methods.

[0137] 2.1 Experimental Setup

[0138] 2.1.1 Dataset

[0139] Each dataset contains two real social networks, and the detailed statistics are shown in Table 1.

[0140] Dataset 1 (Acm-Dblp): This dataset is a collaborator network, where nodes (users) represent authors and edges indicate that two authors have co-authored at least one paper. It is provided by Zhang et al. (Zhang S, Tong H, Jin L, et al. Balancing consistency and disparity in network alignment[C] / / Proceedings of the 27thACM SIGKDD conference on knowledge discovery and data mining. 2021:2212-2222.) and can be obtained online (https: / / github.com / sizhang92 / NextAlign-KDD21).

[0141] Dataset 2 (Online-Offline): Nodes in the online network represent users, and edges represent interactions between users. The offline network is constructed based on the co-occurrence of users in social gatherings. The data can be obtained from (Tang J, Zhang W, Li J, et al. Robust attributed graph alignment via joint structure learning and optimal transport [C] / / IEEE 39th International Conference on Data Engineering. 2023: 1638-1651.) (https: / / github.hscsec.cn / squareRoot3 / SLOTAlign).

[0142] Dataset 3 (Cora1-Cora2): Generate two noise permutations from the citation network Cora to obtain Cora1 and Cora2, insert 10% of the edges into Cora1, remove 15% of the edges from Cora2, and add 10% noise to the node attribute matrix. This dataset can be obtained from (Zhang S, Tong H, Xia Y, et al. Nettrans: Neural cross-network transformation [C] / / Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 2020: 986-996.) (https: / / github.com / yucheny5 / BRIGHT).

[0143] Table 1. Attribute social network dataset information

[0144]

[0145] 2.1.2 Benchmark Method

[0146] In order to evaluate the performance of the method of the present invention, this example compares the method of the present invention with five typical benchmark methods, which are described as follows:

[0147] (1) FINAL (Zhang S, Tong H. Final: Fast attributed network alignment [C] / / Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. 2016: 1345-1354.): This method solves the social network alignment problem based on alignment consistency and uses node / edge attribute information to guide the network alignment process based on topology structure.

[0148] (2) NetTrans (Zhang S, Tong H, Xia Y, et al. Nettrans: Neural cross-network transformation [C] / / Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 2020: 986-996.): This method is based on graph convolutional neural networks and solves the attribute network alignment task from the perspective of network transformation.

[0149] (3) BRIGHT (Yan Y, Zhang S, Tong H. Bright: A bridging algorithm for network alignment [C] / / Proceedings of the web conference. 2021: 3907-3917.): This method uses restarted random walk and graph convolutional neural network to learn user structural feature representation and attribute feature representation respectively, and concatenates the structural features and attribute features as the final features of the user.

[0150] (4)NextAlign(Zhang S, Tong H, Jin L, et al. Balancing consistency and disparity in network alignment[C] / / Proceedings of the 27th ACM SIGKDD conference on knowledge discovery and data mining. 2021: 2212-2222.)0: This method reveals the intrinsic relationship between consistency and disparity in the network alignment problem and studies negative sampling strategies to achieve a balance between consistency and disparity.

[0151] (5) CCNE (Zhang HF, Ren G, Ding X, et al. Collaborative cross-network embedding framework for network alignment [J]. IEEE Transactions on Network Science and Engineering, 2024, 11 (3): 2989-3001.): The CCNE method proposes a collaborative cross-network embedding framework to learn the structural feature representation of nodes by collaboratively optimizing an objective function composed of intra-network and inter-network losses.

[0152] 2.1.3 Evaluation indicators

[0153] The most commonly used precision@k metric is used to evaluate the performance of the proposed method and the baseline method, which is defined as follows:

[0154]

[0155] Where S test represents the test set, |S test | is the number of anchor links in the test set; for anchor links For Node Find the k most similar nodes in the target network (source network) as its top-k candidate set; success i @k means Is it present in In the top-k candidate set, if yes, success i @k=1, otherwise success i @k=0; similarly, success j @k means Is it present in The top-k candidate sets.

[0156] In addition to using the precision@k metric, the Mean Average Precision (MAP) is also used to evaluate model performance, which is defined as follows:

[0157]

[0158] For anchor links The node The similarities with all nodes in the target network (source network) are sorted in descending order as candidate sets. i It means Appear in The position in the candidate set, ra j express Appear in It should be noted that the benchmark method BRIGHT uses a one-way calculation method from the source network to the target network, and this embodiment uniformly changes to a two-way calculation method from the source network to the target network and from the target network to the source network.

[0159] 2.1.4 Experimental parameter settings

[0160] The experiment of this embodiment randomly extracts 20% from the known anchor links as a training set, and the rest as a test set. In order to avoid the influence of experimental randomness and accidental factors on the experimental results, each data set is randomly extracted 5 times, and the experiments are carried out independently in sequence to compare the average values ​​of the experimental results. Specifically, the dimension of the randomly initialized user initial structural feature representation is kept consistent with the node attribute feature dimension. The encoder consists of two layers with dimensions of 300 and 100 respectively, and the decoder dimension is set opposite to the encoder. The number of negative samples K, the learning rate (lr), the batch size (bs), and the threshold margin in the contrast loss are set to 20, 0.01, 512, and 3.0, respectively. Other hyperparameters of the benchmark method are set according to the corresponding article. All experiments in the embodiment are carried out on an AMAX server with a server CPU of Xeon(R)Silver 4210CPU@2.20GHz, GPU is NVIDIACorporation TU102, and memory is 128G.

[0161] 2.2 Experimental Results Analysis

[0162] Table 2. Experimental results on three attribute social network datasets (%)

[0163]

[0164] From Table 2 and Figure 2 It can be seen that:

[0165] (1) Overall, the proposed method outperforms other benchmark methods. Specifically, the precision@1 index of the proposed method on the Acm-Dblp and Online-Offline datasets is improved by 8.8% and 11.8% respectively compared with the best baseline method. On the Cora1-Cora2 dataset, the CLAS and NetTrans methods performed best, among which the CLAS method was slightly better than NetTrans. Compared with other benchmark methods except NetTrans, the CLAS method has significant improvements in both precision@1 and MAP.

[0166] (2) According to the experimental results on three datasets, the NextAlign method performs the worst, which is most obvious on the Online-Offline and Cora1-Cora2 datasets. Taking the Cora1-Cora2 dataset as an example, the precision@1 index of the NextAlign method can only reach 42%, while other benchmark methods exceed 75%. The NextAlign method merges the source network and the target network into a large network based on known anchor links, allowing node attribute information to be transmitted between the source network and the target network, which will introduce more noise and confusion.

[0167] (3) Compared with the BRIGHT method that directly concatenates attribute features and structural features, the accuracy@1 and MAP

[0168] Taking the indicators as an example, the method of the present invention improved by 29.6% and 23.3% on the Acm-Dblp dataset, 14.2% and 12.8% on the Online-Offline dataset, and 14.1% and 10.5% on the Cora1-Cora2 dataset. Although the method of directly splicing attribute and structure information can promote the mutual complementation between the two features, it is impossible to model the deviation between the two. The superiority of the method of the present invention can be seen from the comparison of experimental results.

[0169] (4) Compared with the methods based on GCN fusion attributes and structures, such as NetTrans and CCNE methods, taking precision@1 and MAP indicators as examples, the method of the present invention is 8.8% and 7.2% higher than the NetTrans method on the Acm-Dblp dataset, and 11.8% and 10.9% higher than the CCNE method on the Online-Offline dataset. From the experimental results, although the method based on GCN fusion attributes and structures has demonstrated better network alignment accuracy, it still cannot alleviate the impact of the deviation between the two. The comparison of experimental results also illustrates the effectiveness of consensus constraints in the method of the present invention to a certain extent.

[0170] 2.3 Experimental parameter analysis

[0171] It is used to explore the influence of anchor link ratio r and negative sample number K on attribute social network alignment method. All experiments in this section are conducted on the Online-Offline dataset. When analyzing one of the parameters, the other parameters are kept the same as the overall method of the present invention.

[0172] 2.3.1 Anchor link ratio used for training

[0173] In order to explore the impact of the anchor link ratio in the training set on the experimental results, the anchor link ratio was set to 0.1-0.9, the interval was set to 0.1 each time, and 9 independent experiments were carried out. Figure 3 The average of multiple experimental results with different training set ratios is shown.

[0174] As can be seen from the figure:

[0175] (1) As the proportion of anchor links increases, the performance of all methods continues to improve. However, the experimental results show that the NextAlign method performs the worst and fluctuates, which may also be related to the cross-transmission of node attribute information between the source network and the target network.

[0176] (2) When the anchor link ratio is not higher than 0.2, CCNE is the best among all the benchmark methods. Even when r is only 0.1, it performs better than the proposed method in terms of precision@1 and MAP. However, when the anchor link ratio reaches 0.2, the proposed method shows a very obvious advantage.

[0177] 2.3.2 Impact of the number of negative samples

[0178] In order to explore the impact of the number of negative samples K on the experimental results, the number of negative samples was set to 1, 5, 10, 20, 30, 40, 50 and 60, and 8 independent experiments were carried out. Figure 4 The average value of multiple experimental results under different numbers of negative samples is shown, and it can be seen that:

[0179] (1) For the method of the present invention, the performance reaches the best when the number of negative samples is 5. Although the method of the present invention has some fluctuations with the change of the number of negative samples, it performs better than the baseline method under the conditions of 8 different numbers of negative samples.

[0180] (2) The NextAlign method is most affected by the number of negative samples. For example, when the number of negative samples does not exceed 5, the precision@1 index of the NextAlign method is less than 17%, while the method of the present invention and other benchmark methods still have better performance.

[0181] (3) When the number of negative samples reaches 50 or more, the performance of some methods tends to decline, which may be related to the selection method of negative samples. For example, the present invention constructs negative samples by random sampling, and the benchmark method CCNE selects nodes that are similar to the real corresponding nodes to construct negative samples. This is also an important direction in social network alignment research.

[0182] 2.4 Ablation Experiment

[0183] In order to further demonstrate the effectiveness and rationality of the method of the present invention, this section conducts two sets of ablation experiments from the perspectives of constraint mechanism, attribute consistency and structural consistency.

[0184] 2.4.1 Impact of different constraint mechanisms

[0185] In order to explore the impact of different constraint mechanisms on the method of the present invention, the reconstruction loss constraint, negative sample constraint and consensus constraint are removed from the method of the present invention respectively. The settings of the first group of ablation experiments are as follows:

[0186] Ablation 1: Removing the reconstruction loss constraint. That is, removing the attribute reconstruction loss (3) and the structure reconstruction loss (7) from the overall loss function (13) of the model of the method of the present invention.

[0187] Ablation2: Remove the negative sample constraint. That is, the mean square error loss function obtained only by the positive sample constraint is used instead of the contrast loss (9) obtained by the positive and negative sample constraints.

[0188] Ablation3: Remove the consensus constraint. That is, remove the third term (10) from the overall loss function (13).

[0189] Table 3. Ablation experiment results of exploring constraint mechanism on Online-Offline dataset

[0190]

[0191] The first set of ablation experiments was conducted on the Online-Offline dataset. The hyperparameter settings were consistent with the overall method. The experimental results are shown in Table 3. From the experimental results, we can see that:

[0192] (1) By comparing with Ablation1, it can be found that after removing the reconstruction loss constraint, the performance drops most significantly, with a decrease of 18.0% and 18.7% in precision@1 and MAP indicators, respectively. This shows the key role of the reconstruction process in the attribute social network alignment problem.

[0193] (2) Compared with Ablation2, when the constraint of negative samples is removed, the performance of the method also decreases to a certain extent, with a decrease of 11.0% and 10.9% in precision@1 and MAP indicators respectively. This shows that the introduction of negative samples has a certain positive effect on improving the distinguishability of user feature representation and improving the accuracy of network alignment.

[0194] (3) Compared with Ablation3, after removing the consensus constraint, the precision@1 and MAP indicators of the proposed method are reduced by 4.9% and 5.4% respectively. It can be seen that the consensus constraint also shows certain effectiveness in the attribute social network alignment problem.

[0195] 2.4.2 Impact of attribute consistency and structural consistency

[0196] In order to study the impact of attribute consistency and structural consistency on the attribute social network alignment problem, this section conducts experiments based only on attributes and only on structures, denoted as CLAS-A and CLAS-S respectively. At the same time, experiments on the BRIGHT and CCNE methods based only on relationship structures on three datasets are also added. The second set of ablation experiments is carried out on three datasets, and the experimental results are shown in Table 4.

[0197] Table 4. Effect of property consistency and structural consistency on experimental results (%)

[0198]

[0199] Through the analysis of the experimental results, we can see that:

[0200] (1) Comparison of experimental results of three methods based only on social relationship structure on three datasets. As can be seen from Table 4, the CLAS-S method performs better than BRIGHT and CCNE on all three datasets.

[0201] (2) There are significant differences in attribute consistency and structural consistency among different datasets. The Cora1-Cora2 dataset has stronger attribute consistency. The precision@1 index of the CLAS-A method based only on attribute consistency reaches 98%, while the precision@1 is only about 46% when only the relational structure is used. It can be seen that there is a significant deviation between attributes and structures. However, the structural consistency of the Online-Offline dataset is stronger. The CLAS method applies dual consistency constraints of attributes and structures, and designs consensus constraints to alleviate the impact of the deviation between the two, achieving excellent performance on all three datasets.

[0202] (3) From the experimental results, it can be observed that the precision@1 based on attribute information alone has reached more than 90% on the Cora1-Cora2 dataset. However, some benchmark methods, such as BRIGHT and CCNE methods, fuse attribute information and structural characteristics, which leads to a decrease in method performance. From this phenomenon, it can be seen that the deviation between attributes and structures does affect the determination of the correspondence between users across networks, which further illustrates the effectiveness and rationality of the method of the present invention.

[0203] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An attribute social network alignment method based on attribute and structure consensus learning, characterized by: The following steps are involved: Step 1: Based on the autoencoder, learn the attribute feature representation of the attribute social network users and establish the attribute reconstruction loss; Step 2: Extract the structural features of attribute social network users based on the graph autoencoder and establish the structural reconstruction loss; Step 3: Establish structural consistency constraints and attribute consistency constraints; Step 4: Establish consensus constraints between attribute feature representation and structural feature representation; Step 5: Establish the optimization goal of the method based on the attribute reconstruction loss constraint, the structure reconstruction loss constraint, the attribute consistency constraint, the structure consistency constraint and the consensus constraint, and achieve convergence through iterative training; Step 6: Network alignment: After the method converges, the attribute feature representation and the structural feature representation are concatenated as the final representation of the user, and the network alignment problem is transformed into a ranking retrieval problem based on the similarity measurement.

2. The attribute social network alignment method based on attribute and structure consensus learning according to claim 1 is characterized in that: In step 1, the autoencoder is an autoencoder based on a multi-layer perceptron model.

3. The attribute social network alignment method based on attribute and structure consensus learning according to claim 1 is characterized in that: Step 1 includes: First, a multi-layer perceptron model is used as an encoder to extract user attribute feature representation Z U : Z U =encoder U (X U ) (1) Use another multilayer perceptron model as a decoder to reconstruct the original attributes: Then, based on the mean square error loss function, the attribute reconstruction loss is established by minimizing the difference between the reconstructed attribute and the original attribute. Among them, X represents the attribute feature matrix, U represents the attribute feature, and X U represents the initial attribute feature representation, Represents the reconstructed attribute feature representation, represents the attribute reconstruction loss, ||·|| F represents the Frobenius norm.

4. The attribute social network alignment method based on attribute and structure consensus learning according to claim 1, characterized in that: In step 2, the graph autoencoder uses a graph convolutional neural network as an encoder, and the structure reconstruction loss is established by minimizing the difference between the original adjacency matrix and the reconstructed adjacency matrix.

5. The attribute social network alignment method based on attribute and structure consensus learning according to claim 1, characterized in that: Step 2 includes: Step 2.1: Structural feature encoding, using graph convolutional neural network as the local structural feature encoder, with the adjacency matrix A of the source network s and the adjacency matrix A of the target network t As input, the user information is transferred and aggregated through convolution operation: Among them, A represents the adjacency matrix of the attribute social network, s and t represent the source network and the target network, and H (l-1) and H (l) Respectively represent the input and output of the lth layer GCN, l represents the current GCN layer, Indicates the addition of a self-connected adjacency matrix, I N is the unit matrix, N represents the number of users in the social network, express The degree matrix of represents the weight matrix of the lth layer, σ(·) is the activation function; therefore, the user structure feature representation Z can be obtained V : Z V =encoder V (X V ,A) (5) Among them, V represents the structural characteristics, X V represents the initial structural feature representation, obtained by random initialization; Step 2.2: Structural reconstruction loss. Use the inner product decoder to reconstruct the relationship structure between the source network and the target network. The structural reconstruction loss is established by minimizing the difference between the reconstructed adjacency matrix and the original adjacency matrix. The reconstructed adjacency matrix is ​​expressed as follows: in, To reconstruct the adjacency matrix, Represents Z V is the transpose of , σ(·) is the activation function; the structure reconstruction loss based on the cross entropy loss function is: in, represents the structural reconstruction loss, y and Represent the adjacency matrix A and the reconstructed adjacency matrix respectively. Elements in The reconstruction loss constraints corresponding to step 1 and step 2 are as follows: in, Represents the sum of the two social network attribute reconstruction losses and the structure reconstruction loss.

6. The attribute social network alignment method based on attribute and structure consensus learning according to claim 1, characterized in that: Step 3 is: taking the known anchor links as natural positive samples, constructing negative samples in a random manner, and establishing structural consistency constraints and attribute consistency constraints based on the distance relationship between positive and negative samples.

7. The attribute social network alignment method based on attribute and structure consensus learning according to claim 1, characterized in that: Step 3 includes: Step 3.1: Construct positive and negative samples and link anchors with known correspondences As a natural positive sample, the anchor node v is randomly i s Find K users in the target network Thus construct K negative samples Anchor Node Find K users in the source network Thus, K negative samples are constructed again The set of positive samples and negative samples is called the training sample set, denoted by Among them, ν represents the set of users in the social network, and i, j, m, and n are used to distinguish different users; represents the set of known anchor links between two networks, Represents the set of positive and negative samples used for model training; Step 3.2: Establish contrast loss. For any sample in the training sample set The contrast loss is expressed as follows: in, represents contrast loss, y′ represents sample The label of the positive sample is 1 and the negative sample is 0. Representation sample The Euclidean distance of Indicates the number of samples in the training sample set; margin is the set threshold; The corresponding constraints for step 3 are as follows: in, Represents a consistency constraint.

8. The attribute social network alignment method based on attribute and structure consensus learning according to claim 1, characterized in that: Step 4 includes: First, L2 regularization is used to regularize the attribute feature representation and the structural feature representation, and then the similarity matrix between users under different feature views can be obtained: Among them, S U represents the attribute feature similarity matrix, S V represents the structural feature similarity matrix, Z Unor and Z Vnor Z U and Z V L2 regularization of ; Then, based on the attribute feature similarity matrix S U and the structural feature similarity matrix S V Similar assumptions are made to establish consensus constraints based on the mean square error function. The consensus constraints are expressed as follows: in, Represents a consensus constraint.

9. The attribute social network alignment method based on attribute and structure consensus learning according to claim 1, characterized in that: Step 5 includes: Combine formulas (8), (10) and (12) to establish the overall optimization goal, and use the Adam optimizer to minimize the overall loss function: in, represents the overall loss function; The error back-propagation method is applied to realize the iterative optimization of the proposed method.

10. The attribute social network alignment method based on attribute and structure consensus learning according to claim 1, characterized in that: Step 6 includes: After the training of the method converges, the attribute feature representation and the structural feature representation are concatenated as the final feature representation of the user: Based on similarity measurement, the network alignment problem is transformed into a ranking retrieval problem; for users in the source network or the target network Based on the Euclidean distance, the most similar k users are found in the target network or source network as its top-k candidate set.

Citation Information

Cited By

  • Uncertain adaptive structure-attribute learning graph autoencoder modeling method and system

    CN121683870A