User identity association method for relieving cross-platform heterogeneity

Through the combination of semantic perception and topological perception representation alignment modules, dynamic weighting and adaptive integration of text attributes and topological structure information, the challenges of cross-platform heterogeneity and multi-source information integration are solved, and efficient cross-platform association and accurate link prediction of user identity are achieved.

CN120030505AActive Publication Date: 2025-05-23TIANJIN UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510070949.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-23
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The prior art has challenges in cross-platform heterogeneity and multi-source information integration, and it is difficult to effectively deal with the differences in behavior patterns, differences in social connection patterns, and heterogeneity of text attributes and topological information on different platforms.

Method used

The semantic-aware representation alignment module and topological-aware representation alignment module are adopted, and the adaptive integration mechanism of dynamic weighted text attribute representation and topological attribute representation is combined. The feature representation ability is improved through a consistently enhanced learning mechanism, and an adaptive representation integration module is introduced to generate the final representation to realize cross-platform association of user identities.

Benefits of technology

It significantly reduces the impact of cross-platform heterogeneity on association accuracy, improves the accuracy of cross-network link prediction, and can effectively respond to the challenges brought about by attribute and structural heterogeneity in social networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030505A_ABST
    Figure CN120030505A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of data mining and association rules in an information system, and provides a user identity association method for relieving cross-platform heterogeneity, which effectively relieves text attribute heterogeneity by designing a semantic perception representation alignment module and establishing consistency between user generated contents of different platforms; and the topology structure heterogeneity between networks is reduced by adopting a network densification mechanism through a topology perception representation alignment module. The two modules are trained based on an optimization strategy designed by the invention, and the uniqueness of the neighborhood is kept while the similarity of the anchor nodes is improved to the greatest extent. According to the invention, an adaptive representation integration module is also introduced to dynamically calibrate the weight contribution of text attributes and topological structure information in user representation learning, so that cross-platform user identity association is efficiently realized. The method can be used as a basic module to provide accurate user identification and positioning support for a downstream recommendation system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data mining and association rules in information systems, and discloses a user identity association method for alleviating cross-platform heterogeneity. Background Art

[0002] The field of User Identity Linkage is developing rapidly, and its core goal is to match user accounts on different social platforms. Due to the wide value of this research in practical applications, this field has received much attention and research. Contemporary methods mainly model user identity association as a cross-network anchor node prediction problem by analyzing the similarity of text attributes (such as user-generated content) or the relevance of social relationships, that is, identifying the same user (anchor node pairs) in different social networks. However, these methods face severe challenges in coping with cross-platform heterogeneity and effectively integrating multiple information sources, especially when users exhibit specific attribute information and social interaction characteristics on different platforms.

[0003] Existing models face major challenges in the following three aspects:

[0004] 1. How to solve the problem of cross-network text attribute heterogeneity? The behavior patterns of the same user on different platforms often vary significantly. For example, a researcher may mainly publish papers related to natural language processing at top natural language processing conferences (such as EMNLP), while focusing on computer vision at computer vision conferences (such as ECCV), resulting in significant semantic differences. Current methods usually directly calculate similarity based on text attribute representations and are difficult to effectively cope with this cross-network attribute heterogeneity.

[0005] 2. How to handle cross-network structural heterogeneity? The social connection patterns of users on different platforms often vary greatly. For example, a user's Google Scholar network reflects their academic cooperation relationships, while their connections on LinkedIn represent professional relationships. Existing methods usually independently learn the representations of social network topological attributes and are difficult to capture cross-network structural differences and their impact on user identity association.

[0006] 3. How to effectively integrate text attribute information and topological structure information? User text attributes describe individual characteristics, while network structure reveals social interaction patterns, and both have different importance in user identity association. Many current methods rely on simple concatenation of text attribute information and topological structure information, ignoring the different contribution degrees of these heterogeneous information in user identity association.

[0007] In response to the above challenges, new solutions are urgently needed to better cope with cross-platform heterogeneity and make full use of multi-source information to promote the development of the user identity association field. Summary of the Invention

[0008] In view of the above status quo and its shortcomings, the present invention proposes a user identity association method to alleviate cross-platform heterogeneity. The method aims to provide a framework solution to effectively balance the importance of user text attributes and network topology characteristics, while significantly reducing the impact of cross-platform heterogeneity on association accuracy.

[0009] The technical solution of the present invention is a user identity association method that alleviates cross-platform heterogeneity. Through a semantic-aware representation alignment module and a topology-aware representation alignment module, combined with an adaptive integration mechanism of dynamic weighted text attribute representation and topology attribute representation, cross-platform association of user identity is achieved. Specifically, by designing a semantic-aware representation alignment module (Semantic-aware Representation Alignment) and a topology-aware representation alignment module (Topology-aware Representation Alignment), the consistency-enhanced learning mechanism (Consistency-enhancedLearning) is used to improve the feature representation capability. Finally, the present invention introduces an adaptive representation integration module (AdaptiveRepresentation Incorporation) to generate the final representation, and performs anchor node link prediction (Anchor LinkPrediction), thereby efficiently realizing cross-platform user identity association. The present invention can provide accurate user identification and positioning support for downstream recommendation systems.

[0010] The specific implementation steps are as follows:

[0011] S1 Node Representation Learning: This paper uses Google’s classic vector generation tool (Word2Vec) and large-scale information network embedding method (LINE) to generate initial representations for user-generated text attributes and social network topological attributes, representing the text attribute representation of known anchor nodes. and social network topology attribute representation

[0012] S2 Consistency Enhanced Learning: Through the semantic-aware representation alignment module and the topology-aware representation alignment module, the extracted text attribute representation and topology attribute representation are aligned to enhance the representation consistency of the same user in different networks, and finally generate two enhanced representations.

[0013] The detailed steps of S2 are as follows:

[0014] S2-1 inputs the extracted text attribute representation into the semantic-aware representation alignment module. According to the anchor nodes in the alignment problem, each anchor node is aligned with (i.e., the representation of the same user in different networks) is regarded as a positive sample, and the text attribute representation of the neighboring nodes is regarded as a negative sample to enhance the model's discrimination ability. attribute Optimize and finally generate enhanced text attribute representation.

[0015] S2-2 supplements the neighborhood connections of anchor nodes identified as the same user in different networks through a bidirectional densification mechanism to reduce structural gaps while retaining the inherent characteristics of user relationships. Subsequently, the densified topological attribute representation is input into the topology-aware representation alignment module, using the same positive and negative sample selection strategy as S2-1, and the loss function L is used to select the best topological attribute representation. topology The optimization results in an enhanced representation of topological attributes. The positive and negative sample selection strategies of the semantic perception and topology perception modules are consistent, but the parameters are independent of each other.

[0016] The S2-3 semantic-aware representation alignment module aims to distinguish different users while retaining the specific attribute characteristics of the same user; the topology-aware representation alignment module ensures that nodes maintain similar structural patterns in different networks.

[0017] The total loss function L for joint training of two modules total =λ attribute ·L attribute +λ topology ·L topology , where L attribute and L topology are the loss functions of the semantic-aware representation alignment module and the topology-aware representation alignment module, λ attribute and λ topology is the weighted coefficient of the two loss functions. The total loss function combines the optimization objectives of semantic and structural features.

[0018] S3 Adaptive Representation Integration: Although S2-1 and S2-2 modules achieve effective user identity association by retaining semantic and structural similarities, their independent processing of attribute and structural information fails to fully tap the complementarity of the two. The present invention introduces an adaptive representation integration module to dynamically integrate text attributes and topological structure information to improve user identity association performance.

[0019] This module consists of two key parts: multiple expert networks and router mechanisms. Each expert network is a feedforward neural network that independently processes the input text attribute representation and topological attribute representation. The router dynamically calculates weights based on the input representation to determine the relevance of each expert network. The enhanced text content representation and topological content representation obtained in S2-1 and S2-2 are input into the adaptive representation integration module. After dynamic weighting, the expert network and router will output a weighted expert output node representation. The training loss of the adaptive representation integration model can be expressed as in and They are the anchor node pairs in different networks. The final node representation. ARI When it is the smallest, the distances between pairs of anchor nodes in different networks are the shortest, thus obtaining a further optimized anchor node correlation representation.

[0020] S4 Anchor node link prediction: After obtaining the unified representation, this module identifies users across networks through distance calculation. Specifically, for the anchor node pairs in the two networks obtained from S3 (the final representation values ​​are and ), the module will calculate their Euclidean distance as The smaller the Euclidean distance, the greater the possibility that the two nodes represent the same natural person.

[0021] Beneficial Effects

[0022] The present invention mitigates the heterogeneity of user-generated content (text attributes) and social networks (topological attributes) across different platforms and identifies user identities (i.e., anchor nodes). The present invention provides a framework for linking user identities across networks that can effectively address the challenges posed by attribute and structural heterogeneity in social networks.

[0023] The present invention proposes a semantic-aware representation alignment module, which successfully aligns the representations of user-generated content while retaining specific network features, effectively reducing the impact of different content patterns across networks; and introduces a topology-aware representation alignment module, which significantly enhances structural consistency while maintaining basic topological information through an innovative network densification mechanism.

[0024] In addition, the present invention combines an adaptive representation integration architecture to dynamically integrate multi-view representations, providing a more flexible and efficient solution for representation fusion.

[0025] The present invention significantly improves the accuracy of cross-network link prediction, that is, the accuracy of cross-network user identity association. The present invention can be used as a basic module and is widely used in the development of related services such as cross-network recommendation, user portrait construction, and malicious entity detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 The model structure is trained for the present invention.

[0027] Figure 2 Comparison chart of experimental results. DETAILED DESCRIPTION

[0028] The present invention will be further described below in conjunction with the accompanying drawings.

[0029] S1 Node representation learning: The present invention performs feature representation learning for user-generated text attributes and social network topological attributes. The present invention uses Google's classic vector generation tool Word2Vec model to perform representation learning on user-generated text information (such as published content, introduction or related text attributes). Figure 1 As shown in the figure, the text information is first segmented and decomposed into multiple words (word 1 to word N). Then, each word is mapped to a low-dimensional embedding vector through Word2Vec, and the average of multiple word embedding vectors is taken for aggregation, and finally the text representation is generated through the Softmax classifier. For the social network topology attributes of the anchor node, the large-scale information network embedding method (LINE) is used to generate a low-dimensional structure representation vector of the node based on the first-order similarity and second-order similarity of the node.

[0030] Learning of S2 consistency enhancement: The present invention inputs the text attribute representation extracted by S1 into the semantic-aware representation alignment module, and inputs the topological attribute representation of the node into the topological-aware representation alignment module. Through the semantic-aware representation alignment module and the topological-aware representation alignment module, the extracted text attribute representation and topological attribute representation are aligned to enhance the representation consistency of the same user in different networks, and finally generate two enhanced representations.

[0031] The detailed steps of S2 are as follows:

[0032] S2-1 inputs the text attribute representation into the semantic perception representation alignment module. Extract text attribute representation from the anchor node pairs (Train Anchors) of the source network and the target network. (i.e., the representation of the same user in different networks) as positive samples (anchors embeds1 and anchor embeds2 ), and select text attribute representations from neighboring nodes as negative samples (neg embeds1 and neg embeds2 ), improve the model's ability to distinguish the consistency of positive samples and the discrimination of negative samples. Through the module's loss function L attribute Optimize and finally generate enhanced text attribute representation.

[0033] The S2-2 topology-aware representation alignment module is similar to the semantic-aware module, but it performs bidirectional densification on the topological structure of the graph before processing. Figure 1 As shown, given the anchor node A in the source network 1 and the corresponding anchor node B in the target network 1 , if A 1 There is a neighbor A 2 , which corresponds to anchor node B 2 exists in the target network, but is not related to B 1No connection, just B 1 and B 2 Then, the topological attribute representation is used as input to perform batch division of positive and negative samples through GPU acceleration.

[0034] S2-3 uses the contrastive loss function (contrastive_loss) to calculate the loss of positive and negative sample pairs, and accumulates the loss of each batch into the total loss (combined_loss) to optimize the model's learning of the consistency of positive samples and the discrimination of negative samples. topology The optimization results in an enhanced representation of topological attributes. The positive and negative sample selection strategies of the semantic perception and topology perception modules are consistent, but the parameters are independent of each other.

[0035] The present invention realizes joint training by weighted accumulation of the loss functions of the semantic perception and topology perception modules. The total loss function is expressed as: L total =λ attribute ·L attribute +λ structure ·L topology .

[0036] S3 Adaptive representation integration: In order to further improve the effect of representation fusion, the present invention introduces an adaptive representation integration module to generate a joint representation through dynamic weight allocation and fusion.

[0037] The module consists of two key parts: multiple expert networks and router mechanisms;

[0038] Each expert network is a feed-forward neural network that independently processes the input text attribute representation and topological attribute representation;

[0039] The router dynamically calculates the weights based on the input representation and determines the relevance of each expert network;

[0040] The enhanced text content representation and topology content representation obtained by S2 are input into the adaptive representation integration module, and the expert network and router output a weighted expert output node representation after dynamic weighting;

[0041] The training loss of the adaptive representation integration model is expressed as in and They are the anchor node pairs in different networks. The final node representation, when L ARI When it is the smallest, the distances between pairs of anchor nodes in different networks are the shortest, thus obtaining a further optimized anchor node correlation representation;

[0042] Each expert network consists of two fully connected layers (nn.Linear) and a nonlinear activation function (nn.ReLU) for text attribute representation or topological attribute representation. For the router mechanism (Router), the present invention uses a fully connected layer to represent the text attribute (x a ) and topological properties characterization (x s ) respectively generate expert weights, and use the softmax classifier to normalize them to ensure that the sum of the weights is 1. The multi-expert fusion layer (Mixture of Experts Layer) contains multiple expert networks (implemented by nn.ModuleList), which calculates the output of all experts for the input text attribute representation and topological attribute representation, and uses the weights generated by the router (weights a and weights s ) is weighted and summed to generate the final text attribute representation and topological attribute representation output. The weighted text attribute representation output and topological attribute representation output are concatenated through torch.cat to form a joint representation. The adaptive layer (Mixture of Experts Sequential Layer) consists of multiple multi-expert fusion layers, each of which includes input normalization (nn.LayerNorm), multi-expert fusion (Mixture of ExpertsLayer), output normalization (nn.LayerNorm) and residual connection (conbined input +MoE output ), improving training stability and generating the final joint representation.

[0043] S4 anchor node link prediction evaluates model performance through Top-K accuracy and mean reciprocal rank (MRR). The present invention measures the strength of association between nodes by calculating the Euclidean distance (or other similarity indicators) between anchor node embeddings. All possible node pairs are sorted according to the calculation results. The smaller the distance, the stronger the predicted association. The Top-K accuracy statistics test whether the actual node pairs (positive samples) in the test set appear in the top K node pairs predicted, reflecting the accuracy of the model in local prediction. The mean reciprocal rank (MRR) calculates the reciprocal rank (1 / Rank) of each positive sample in the test set in the predicted sorting, and takes the average of the reciprocal ranks of all positive samples to obtain MRR. The MRR indicator reflects the overall ranking performance of the positive sample in the prediction results. The higher the value, the faster and more accurately the model can identify the positive sample. Through the dual evaluation of Top-K accuracy and MRR, the performance of the model in local prediction accuracy (Top-K) and global ranking quality (MRR) can be comprehensively measured, ensuring that the model can not only effectively identify strongly associated node pairs, but also reasonably rank the association strength of all node pairs.

[0044] Finally, the present invention enhances the consistency of user identity association across networks by integrating user text attribute representation and topological attribute representation, improves feature expression capability through adaptive fusion, and significantly improves the accuracy of user identity association prediction.

[0045] Finally, the cross-network link prediction scores generated by the present invention indicate the relevance of cross-network user identities. Figure 2As shown in the figure, it shows that the present invention (i.e., EFC-UIL in the figure) has the best performance on three datasets compared with five baseline methods. DBLP_1 and DBLP_2 are public datasets from the DBLP database (a database of journals and conference papers in the field of computer science), and WD is the Weibo Douban dataset. The five baseline methods are: MAUIL is a semi-supervised framework that integrates multi-granular text features (character level, word level and topic level) with structural information of user alignment; GAlign framework aligns networks through multi-order node representations learned by graph convolutional networks, and combines data enhancement mechanism to deal with noise; CENALP is a unified framework that uses cross-network Skip-gram model to capture network topology representation, which can jointly optimize network alignment and link prediction; Grad-Align is a graph neural network method that uses Tversky index to measure asymmetric set similarity between node representations to achieve user identity alignment; NeXtAlign is a method that adjusts the relative position of anchor nodes after each iteration through a restarted random walk method in an attention-based architecture to incorporate relative position information of anchor nodes; NetTrans is based on an encoder-decoder architecture that can learn nonlinear transformations to capture structural and attribute information between networks for alignment. The prediction results can be applied to a variety of practical scenarios, such as cross-network recommendation, user portrait construction and malicious entity detection, to provide support for the development of related services.

Claims

1. A user identity association method for alleviating cross-platform heterogeneity, characterized in that: Through the semantic-aware representation alignment module and the topology-aware representation alignment module, combined with the adaptive integration mechanism of dynamic weighted text attribute representation and topology attribute representation, cross-platform association of user identity is achieved, which specifically includes the following steps: S1 Node Representation Learning: Feature learning is performed for user-generated text attributes and social network topology attributes respectively; S2 Consistency Enhanced Learning: Through the semantic-aware representation alignment module and the topology-aware representation alignment module, the extracted text attribute representation and topology attribute representation are aligned to enhance the representation consistency of the same user in different networks, and the enhanced text content representation and topology content representation are obtained; S3 inputs the enhanced text content representation and topology content representation obtained by S2 into the adaptive representation integration model, and the expert network and router output a weighted expert output node representation after dynamic weighting; The training loss of the adaptive representation integration model is expressed as in and They are the anchor node pairs in different networks. The final node representation, when L ARI When it is the smallest, the distances between pairs of anchor nodes in different networks are the shortest, thus obtaining a further optimized anchor node correlation representation; S4 Anchor node link prediction: Identify users across networks through distance calculation. For the anchor node pairs in the two networks obtained from S3, the final representation values ​​are and The module calculates the Euclidean distance as 2. A user identity association method for alleviating cross-platform heterogeneity according to claim 1, characterized in that: The step S1 uses a vector generation tool and a large-scale information network embedding method to generate an initial representation, representing the text attribute representation of the known anchor node. and social network topology attribute representation 3. A user identity association method for alleviating cross-platform heterogeneity according to claim 1, characterized in that: The step S2 is specifically as follows: S2-1 inputs the text attribute representation into the semantic-aware representation alignment module. According to the anchor nodes in the alignment problem, each anchor node pair is regarded as a positive sample, and the text attribute representation of the neighboring node is regarded as a negative sample. attribute Optimize and finally generate enhanced text attribute representation; S2-2 complements the neighborhood connections of anchor nodes confirmed as the same user in different networks through a bidirectional densification mechanism, inputs the densified topological attribute representation into the topology-aware representation alignment module, and uses the same positive and negative sample selection strategy as S2-1. topology Optimize and obtain enhanced topological property representation; The positive and negative sample selection strategies of the semantic perception and topology perception modules are consistent, but their parameters are independent of each other; The S2-3 semantic-aware representation alignment module aims to distinguish different users while retaining the specific attribute characteristics of the same user; the topology-aware representation alignment module ensures that nodes maintain similar structural patterns in different networks; The total loss function L for joint training of the two modules total =λ attribute ·L attribute +λ topology ·L topology , the optimization goal of integrating semantic and structural features, where L attribute and L topology are the loss functions of the semantic-aware representation alignment module and the topology-aware representation alignment module, λ attribute and λ topology It is the weighted coefficient of the two loss functions.

4. A user identity association method for alleviating cross-platform heterogeneity according to claim 1, characterized in that: The step S3 is specifically as follows: the adaptive representation integration model consists of two key parts: multiple expert networks and a router mechanism; Each expert network is a feed-forward neural network that independently processes the input text attribute representation and topological attribute representation; The router dynamically calculates weights based on the input representation and determines the relevance of each expert network.

Citation Information

Patent Citations

  • Cross-social network user identity recognition method and system based on machine learning

    CN109753602A

  • Cross-social network virtual identity association method and device based on multi-modal fusion and expression alignment

    CN115828109A

  • Multi-network identity alignment system and method based on same-space user feature transmission

    CN116628515A

  • User alignment method of cross-social media network based on graph neural network

    CN116681540A

  • Content matching method and device, computer equipment, storage medium and program product

    CN117688390A