Cross-platform processing method, device, equipment and medium for user identity alignment

By constructing a heterogeneous graph neural network model and deep learning technology to encode and fuse user behavior data, the accuracy and efficiency of cross-platform user identity alignment is solved, and more accurate user identity recognition and more efficient social network data processing is achieved.

CN119515584BActive Publication Date: 2025-07-25BEIJING ZHONGWANG ZHICE TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411379519.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2025-07-25
Estimated Expiration
2044-09-30

Smart Images

  • Figure CN119515584B_ABST
    Figure CN119515584B_ABST
Patent Text Reader

Abstract

The present application discloses a processing method, device, equipment and medium for cross-platform user identity alignment, including: obtaining user behavior data of multiple users on different platforms; determining identity information corresponding to the user behavior data according to a pre-trained identity alignment model, wherein the pre-trained identity alignment model is obtained by using sample user data and an initial training model on different platforms to obtain embedding vectors of heterogeneous information, encoding the embedding vectors of the heterogeneous information to obtain an encoding result, and performing adaptive weight aggregation on the encoding result; determining alignment information of the multiple users according to the identity information, and by constructing a heterogeneous graph neural network model, comprehensively considering various heterogeneous information of users and using deep learning technology to encode and fuse this information, so as to achieve more accurate and efficient cross-platform user identity alignment for user identity alignment processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and particularly relates to a cross-platform processing method, device, equipment, and medium for user identity alignment. Background Art

[0002] With the rapid development of social network technology and the diversification of online social platforms, the identity information of users on different social network platforms has shown explosive growth. This information contains rich user behavior and social interaction data, which is of great value for aspects such as user behavior analysis, personalized content recommendation, and network security management. However, due to factors such as data isolation between different platforms and user privacy protection, it is often difficult to obtain the identity data of users in different networks, which greatly limits the comprehensive understanding and effective utilization of the cross-platform behavior patterns of users.

[0003] Most of the existing user identity alignment methods are limited to single-dimensional attribute matching analysis or local network structure mining, and fail to fully utilize the multi-dimensional characteristics shown by users in social networks and the complex connections between user attributes, content, and social relationships. At the same time, traditional machine learning methods often have difficulty effectively capturing deep user characteristics and global social contexts when dealing with such high-dimensional, highly heterogeneous data containing rich social information. How to comprehensively process deep user characteristics and global social contexts and then identify cross-platform user identities is an urgent problem to be solved currently. Summary of the Invention

[0004] This application aims to provide a cross-platform processing method, device, equipment, and medium for user identity alignment to solve the deficiencies in the prior art. The technical problems to be solved by this application are achieved through the following technical solutions.

[0005] In a first aspect, an embodiment of this application provides a cross-platform processing method for user identity alignment. The method includes:

[0006] Obtain user behavior data of multiple users on different platforms;

[0007] According to a pre-trained identity alignment model, determine the identity information corresponding to the user behavior data; wherein, the pre-trained identity alignment model is obtained by using sample user data and an initial training model on different platforms to obtain embedded vectors of heterogeneous information, encoding the embedded vectors of the heterogeneous information to obtain an encoding result, and performing adaptive weight aggregation on the encoding result;

[0008] According to the identity information, determine the alignment information of the multiple users.

[0009] Optionally, the identity alignment model is obtained in the following manner:

[0010] Obtain sample user data from different platforms;

[0011] Determine a user heterogeneous information dataset based on the sample user data;

[0012] Encode the heterogeneous information in the user heterogeneous information dataset based on a deep learning network model to obtain an embedding vector of the heterogeneous information, where the embedding vector of the heterogeneous information at least includes an embedding vector of text attributes, the content embedding vector, and the time information embedding vector;

[0013] Construct a heterogeneous graph based on the embedding vector of the heterogeneous information;

[0014] Perform adaptive weight aggregation on the information nodes of each user heterogeneous type in the heterogeneous graph to obtain a model input vector;

[0015] Input the model input vectors of users from different platforms into a multi-layer perceptron model respectively to obtain output results corresponding to the model input vectors of users from different platforms;

[0016] Determine whether the sample user data from different platforms comes from the same user according to the output results corresponding to the model input vectors of users from different platforms.

[0017] Optionally, the encoding of the heterogeneous information in the user heterogeneous information dataset based on a deep learning network model to obtain an embedding vector of the heterogeneous information includes:

[0018] Use a BERT model to encode the basic attribute information in the sample user data to obtain an embedding vector of text attributes; where the basic attribute information at least includes a username and a nickname;

[0019] Use a BERTopic model to encode the user-generated content in the sample user data to obtain a content embedding vector, where the content embedding vector at least includes a topic information embedding vector and an interest information embedding vector; encode the timestamp information in the sample user data to obtain a time information embedding vector;

[0020] Fuse the embedding vector of the text attributes, the content embedding vector, and the time information embedding vector to obtain a multi-dimensional embedding vector of the heterogeneous information.

[0021] Optionally, the construction of the heterogeneous graph based on the embedding vector of the heterogeneous information includes:

[0022] Taking the embedding vectors of the text attributes, the content embedding vector, and the time information embedding vector as nodes, and taking the relationship between the social relationship and the content relationship as edges, the heterogeneous graph is determined.

[0023] Optionally, the adaptive weight aggregation of the information nodes of each user heterogeneous type in the heterogeneous graph to obtain a model input vector includes:

[0024] Based on the graph attention model, corresponding weight values are set for the nodes of the embedding vectors of each text attribute, the content embedding vector, and the time information embedding vector respectively;

[0025] According to the embedding vector and the corresponding weight value, the model input vector is determined.

[0026] In a second aspect, an embodiment of the present application provides a cross-platform user identity alignment processing device, and the device includes:

[0027] An acquisition module, configured to acquire user behavior data of multiple users on different platforms;

[0028] A determination module, configured to determine identity information corresponding to the user behavior data according to a pre-trained identity alignment model; wherein, the pre-trained identity alignment model is obtained by using sample user data and an initial training model on different platforms to obtain embedding vectors of heterogeneous information, encoding the embedding vectors of the heterogeneous information to obtain an encoding result, and performing adaptive weight aggregation on the encoding result;

[0029] An alignment module, configured to determine alignment information of the multiple users according to the identity information.

[0030] Optionally, the device further includes a training module, and the training module is configured to:

[0031] Acquire sample user data on different platforms;

[0032] Determine a user heterogeneous information data set according to the sample user data;

[0033] Encoding the heterogeneous information in the user heterogeneous information data set based on a deep learning network model to obtain embedding vectors of the heterogeneous information, where the embedding vectors of the heterogeneous information at least include embedding vectors of text attributes, the content embedding vector, and the time information embedding vector;

[0034] Construct a heterogeneous graph according to the embedding vectors of the heterogeneous information;

[0035] Performing adaptive weight aggregation on the information nodes of each user heterogeneous type in the heterogeneous graph to obtain a model input vector;

[0036] Input the model input vectors of users on different platforms into a multi-layer perceptron model respectively to obtain output results corresponding to the model input vectors of users on different platforms;

[0037] Determine whether the sample user data on different platforms comes from the same user according to the output results corresponding to the model input vectors of users on different platforms.

[0038] Optionally, the training module is used for:

[0039] Use the BERT model to encode the basic attribute information in the sample user data to obtain embedded vectors of text attributes; wherein, the basic attribute information at least includes the user name and nickname;

[0040] Use the BERTopic model to encode the user-generated content in the sample user data to obtain content embedded vectors, wherein the content embedded vectors at least include topic information embedded vectors and interest information embedded vectors; encode the timestamp information in the sample user data to obtain time information embedded vectors;

[0041] Perform fusion processing on the embedded vectors of the text attributes, the content embedded vectors and the time information embedded vectors to obtain embedded vectors of multi-dimensional heterogeneous information.

[0042] Optionally, the training module is used for:

[0043] Use the embedded vectors of the text attributes, the content embedded vectors and the time information embedded vectors as nodes, and use the relationship between the social relationship and the content relationship as edges to determine the heterogeneous graph.

[0044] Optionally, the training module is used for:

[0045] Based on the graph attention model, set corresponding weight values for the nodes of the embedded vectors of each text attribute, the content embedded vectors and the time information embedded vectors respectively;

[0046] Determine the model input vector according to the embedded vector and the corresponding weight value.

[0047] In a third aspect, an embodiment of the present application provides a terminal device, including: at least one processor and a memory;

[0048] The memory stores a computer program; the at least one processor executes the computer program stored in the memory to implement the cross-platform user identity alignment processing method provided in the first aspect.

[0049] Fourthly, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed, it implements the cross-platform user identity alignment processing method provided in the first aspect.

[0050] The embodiments of the present application include the following advantages:

[0051] The cross-platform user identity alignment processing method, device, equipment and medium provided by the embodiments of the present application obtain the user behavior data of multiple users on different platforms; determine the identity information corresponding to the user behavior data according to a pre-trained identity alignment model, wherein the pre-trained identity alignment model is obtained by using the sample user data and the initial training model on different platforms to obtain the embedded vectors of heterogeneous information, encoding the embedded vectors of the heterogeneous information to obtain an encoding result, and performing adaptive weight aggregation on the encoding result; determine the alignment information of the multiple users according to the identity information. By constructing a heterogeneous graph neural network model, various heterogeneous information of users is comprehensively considered, and these information are encoded and fused by using deep learning technology to achieve more accurate and efficient user identity alignment, which not only improves the accuracy of user identity alignment, but also enhances the adaptability and generalization ability of the model to data of different types of social platforms. Through the automated feature extraction and weight assignment mechanism, the present invention also significantly improves the efficiency of processing large-scale social network data, providing an effective technical means for social network analysis and user identity management. Description of the Drawings

[0052] In order to more clearly illustrate the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0053] Figure 1 It is a flowchart of a cross-platform user identity alignment processing method provided by an embodiment of the present application;

[0054] Figure 2 It is a schematic flowchart of a cross-platform user identity alignment method based on a heterogeneous graph neural network provided by an embodiment of the present application;

[0055] Figure 3 It is a schematic flowchart of cross-platform user attribute embedding extraction based on a heterogeneous graph neural network provided by an embodiment of the present invention;

[0056] Figure 4Schematic flow chart of cross - platform user attribute enhancement and unified embedding representation provided by an embodiment of the present invention;

[0057] Figure 5 Block diagram of the structure of an embodiment of a cross - platform user identity alignment processing device of the present application;

[0058] Figure 6 Schematic diagram of the structure of a terminal device of the present application. Detailed implementation manners

[0059] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0060] An embodiment of the present application provides a cross - platform user identity alignment processing method for aligning user identities. The execution subject of this embodiment is a cross - platform user identity alignment processing device, which is set on a terminal device. For example, the terminal device at least includes a computer terminal, etc.

[0061] Referring to Figure 1 , a step flow chart of an embodiment of a cross - platform user identity alignment processing method of the present application is shown. The method may specifically include the following steps:

[0062] S101. Obtain user behavior data of multiple users on different platforms;

[0063] Specifically, the terminal device obtains user behavior data of different platforms through each interface.

[0064] S102. Determine the identity information corresponding to the user behavior data according to a pre - trained identity alignment model; among them, the pre - trained identity alignment model is obtained by using sample user data of different platforms and an initial training model, obtaining an embedded vector of heterogeneous information, encoding the embedded vector of heterogeneous information to obtain an encoding result, and performing adaptive weight aggregation on the encoding result;

[0065] Specifically, the terminal device obtains sample user data of different platforms, inputs the sample user data into the initial training model to obtain an embedded vector of heterogeneous information, encodes the embedded vector of heterogeneous information to obtain an encoding result, performs adaptive weight aggregation on the encoding result, and continuously trains the initial training model to obtain an identity alignment model, which is used to determine whether user data on different platforms comes from the same user.

[0066] S103. Determine the alignment information of multiple users according to the identity information.

[0067] In the embodiments of the present application, deep feature information between nodes and nodes is captured based on BERTopic and BiLSTM. It can not only capture the complex relationships between users on different social platforms, but also learn the internal connections between user attributes, content, and social behaviors. Thus, the accuracy and efficiency of user identity alignment are significantly improved. By leveraging the advantages of GNN and BERT, the problem of user identity alignment is effectively solved, the ability to capture the cross-platform behavior characteristics of users is enhanced, and the precise alignment of user identities is successfully achieved. This not only improves the accuracy of user identity recognition, but also better understands and mines the complex behavior patterns of users in the social network. Therefore, in practical applications, such as personalized recommendation, advertising targeting, network security and other fields, the personalization level and security of services are significantly improved.

[0068] The cross-platform user identity alignment processing method provided by the embodiments of the present application includes: obtaining user behavior data of multiple users on different platforms; determining identity information corresponding to the user behavior data according to a pre-trained identity alignment model, where the pre-trained identity alignment model is obtained by using sample user data and an initial training model on different platforms to obtain embedded vectors of heterogeneous information, encoding the embedded vectors of heterogeneous information to obtain an encoding result, and performing adaptive weight aggregation on the encoding result; determining the alignment information of multiple users according to the identity information. By constructing a heterogeneous graph neural network model, various heterogeneous information of users is comprehensively considered, and these information are encoded and fused by using deep learning technology to achieve more accurate and efficient user identity alignment. This not only improves the accuracy of user identity alignment, but also enhances the adaptability and generalization ability of the model to data of different types of social platforms. Through an automated feature extraction and weight allocation mechanism, the present invention also significantly improves the efficiency of processing large-scale social network data, providing an effective technical means for social network analysis and user identity management.

[0069] Another embodiment of the present application further supplements and explains the cross-platform user identity alignment processing method provided in the above embodiment.

[0070] Figure 2 As shown in Figure 2 the following, it includes:

[0071] S11: Data collection and preprocessing: Collect user data from different platforms through an interface, clean the data, remove noise and standardize the format, construct a user heterogeneous information data set and divide it into a training set and a test set;

[0072] S12: User attribute embedding extraction: Encode various heterogeneous information of the user based on BERT and BiLSTM to extract the embedding representation of the user attributes in a unified manner;

[0073] S13: User attribute enhancement and unified embedding representation: Use the BERTopic model to extract the text style and interests of the user, integrate these deep features into the heterogeneous graph, and perform adaptive weight aggregation on the user heterogeneous type information nodes through GAT (Graph Attention Network) to obtain a unified user embedding representation;

[0074] S14: User identity alignment: Based on the MLP model, input the user embedding representations in two different platforms and predict whether they belong to the same user;

[0075] S15: Model performance evaluation: Evaluate the performance of the model in the user alignment task and use multiple evaluation metrics to verify the alignment accuracy of the model.

[0076] Optionally, the identity alignment model is obtained in the following way:

[0077] Obtain sample user data from different platforms;

[0078] Specifically, collect user data from different platforms through an interface, clean the data, remove noise, and standardize the format, construct a user heterogeneous information dataset and divide it into a training set and a test set;

[0079] Determine the user heterogeneous information dataset according to the sample user data;

[0080] Encode the heterogeneous information in the user heterogeneous information dataset based on a deep learning network model to obtain the embedding vectors of the heterogeneous information, where the embedding vectors of the heterogeneous information at least include the embedding vectors of text attributes, content embedding vectors, and time information embedding vectors;

[0081] Construct a heterogeneous graph according to the embedding vectors of the heterogeneous information;

[0082] Perform adaptive weight aggregation on the information of each user heterogeneous type information node in the heterogeneous graph to obtain the model input vector;

[0083] Input the model input vectors of the users in different platforms into the multi-layer perceptron model respectively to obtain the output results corresponding to the model input vectors of the users in different platforms;

[0084] Determine whether the sample user data from different platforms comes from the same user according to the output results corresponding to the model input vectors of the users in different platforms, receive the user heterogeneous information data from the data acquisition module, and align the user identities using the heterogeneous graph neural network architecture.

[0085] Evaluate and improve the results of user identity alignment using predefined evaluation metrics, and objectively evaluate the effect of user alignment to ensure the accuracy and reliability of the user identity alignment results.

[0086] Optionally, encode the heterogeneous information in the user heterogeneous information dataset based on a deep learning network model to obtain the embedding vectors of the heterogeneous information, including:

[0087] Use the BERT model to encode the basic attribute information in the sample user data to obtain the embedding vectors of the text attributes; among them, the basic attribute information includes at least the username and nickname;

[0088] Use the BERTopic model to encode the user-generated content in the sample user data to obtain the content embedding vectors, where the content embedding vectors include at least the topic information embedding vectors and the interest information embedding vectors; encode the timestamp information in the sample user data to obtain the time information embedding vectors;

[0089] Fuse the embedding vectors of the text attributes, the content embedding vectors, and the time information embedding vectors to obtain the embedding vectors of the multi-dimensional heterogeneous information.

[0090] Specifically, the user attribute embedding extraction step in S12 includes:

[0091] Construct a text attribute embedding module: Encode the basic attribute information of the user (such as username, nickname, personal profile, etc.) using the BERT model to obtain a deep semantic representation;

[0092] Construct a content embedding module: Process the user-generated content (such as text, images), where the text content obtains the embedding representation through the BERT model, and the image content extracts features through the pre-trained CNN model and then uses BERT for further encoding;

[0093] Among them: BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained deep learning model that can achieve good results in various NLP tasks. BERTopic is a text clustering tool based on BERT, mainly used for text clustering and topic classification. It encodes the text, extracts the features in the text, and then clusters the features to group similar texts together. The main advantage of BERTopic is that it uses the BERT model, which has strong capabilities in processing natural language. The BERT model can capture the context information in the text and extract important features from it. In addition, the BERT model can be fine-tuned through multi-task learning to make it more suitable for specific text clustering tasks. In BERTopic, the text is first preprocessed into word-level vector representations. Then, the BERT model is used to encode these vectors to obtain the context embedding representations of each word. These representations are further processed into sentence-level vector representations, and then they are clustered through a clustering algorithm. Finally, by analyzing each cluster, the topic classification result of the text can be obtained.

[0094] Construct a time information embedding module: Standardize the user's timestamp information (such as registration time, activity record time) and convert it into a format suitable for model input;

[0095] Construct a sequential information embedding module: Use a BiLSTM model to process sequential data, such as the user's activity logs or time series data, to capture temporal dependencies;

[0096] Heterogeneous information fusion module: Fuse the embedding representations of the above different types of attributes to form a comprehensive user attribute embedding vector to comprehensively reflect the multi-dimensional features of the user.

[0097] Figure 3 The following is a schematic flowchart of cross-platform user attribute embedding extraction based on a heterogeneous graph neural network provided by an embodiment of the present invention, which specifically includes:

[0098] S21. Construct a text attribute embedding module: Encode the user's basic attribute information (such as username, nickname, personal profile, etc.) using the BERT model to obtain deep semantic representations;

[0099] Considering the co - existence of short - text attributes such as username and nickname and the long - text attribute of tweet text, the model extracts text features according to the MaxPooling method. For short - text attributes, generally one segment is enough, so it is not affected by splitting and the Pooling method. For long texts like tweets, they can be input into the BERT model in a split way, and then the features of multiple segments are aggregated through MaxPooling. As shown in Formula 1, the model first splits the filtered text into texts according to the input length specified by BERT, then sends the set of segmented texts into the BERT pre - trained model for embedding extraction, and then performs MaxPooling on the set of text - segment embedding representations to obtain the final embedding representation e of a text text .

[0100] e text = MaxPooling(BERT(texts)) (Formula 1)

[0101] S22. Construct a content embedding module: Process user - generated content (such as text, images). For text content, obtain the embedding representation through the BERT model. After the image content extracts features through a pre - trained CNN model, use BERT for further encoding;

[0102] S23. Construct a time - information embedding module: Standardize the user's timestamp information (such as registration time, activity record time) and convert it into a format suitable for model input;

[0103] Time is a crucial piece of information in social network platforms, such as the user registration time, the user login time, and the tweet posting time, etc. Such information is a very critical factor in user characteristics. These pieces of information can be used to weight other features or can be processed uniformly as one feature through feature embedding. However, the ways of representing time in different social network platforms are different. Some social network platforms use the Unix timestamp format (such as 1675524522), the ISO 8601 (International Organization for Standardization) format (such as 2023-02-04T15:28:42.178Z), the RFC 1123 (Request For Comments) format (such as Sat, 04 Feb 2023 15:28:42 GMT), etc. It is difficult to process such time with non-uniform formats. The time attributes with non-uniform formats will also lead to mismatches in embeddings between different platforms in the end. Moreover, the BERT pre-trained model can only accept time in text format. Therefore, the model converts all time into a unified text form. First, the model converts time in various different formats into Unix timestamps to obtain a numerical type of time, and then through time formatting operations, converts the numerical type of time into a text time in the yyyy-MM-dd HH:mm:ss format. Then, the time text in the unified format is processed according to the text embedding extraction task process to obtain the embedding representation e of the final time. time 。

[0104] S24. Construct a serialized information embedding module: Use a BiLSTM model to process serialized data, such as user activity logs or time series data, to capture dependencies over time.

[0105] S25. Heterogeneous information fusion module: Fuse the embedding representations of the above different types of attributes to form a comprehensive user attribute embedding vector to comprehensively reflect the multi-dimensional characteristics of the user. Optionally, construct a heterogeneous graph based on the embedding vectors of heterogeneous information, including:

[0106] Use the embedding vectors of text attributes, content embedding vectors, and time information embedding vectors as nodes, and use the relationships between social relationships and content relationships as edges to determine the heterogeneous graph.

[0107] Figure 3 This is a schematic flowchart of a cross-platform user attribute enhancement and unified embedding representation based on a heterogeneous graph neural network provided by an embodiment of the present invention.

[0108] S31. Build a text style and interest extraction module: Use the BERTopic model to perform topic modeling on the text content generated by users, extract the topic information in the text, and convert it into an embedding vector;

[0109] After sending the corpus and parameters into the BERTopic framework, a topic model TopicModel is obtained. This topic model can output a topicText for each input document doc, which is a text type topic.

[0110] After training the topic model TopicModel, the model will first perform topic extraction on each tweet. That is, send all the text in tweet set V c into the topic model topicText to obtain the text type topics TopicModel of all tweets, as shown in Formula 2.

[0111] topicText = topicModel(getText(V c )) (Formula 2)

[0112] To maintain an embedding vector of a unified length of 768, the model uses a topic indexer topicIndexer to convert each text topic topicText into a real number topic topicIndex. Its implementation method is just a mapping function that maps the text to a real number. As shown in Formula 3, each text topic input into the topic indexer will be converted into a real number between [0, 767].

[0113] topicIndex = TopicIndexer(topicText), topicIndex ∈ [0, 767] (Formula 3)

[0114] S32. Build a heterogeneous graph module: Use various user attributes embeddings, style embeddings, and interest embeddings as nodes, and build edges between nodes through social relationships and content relationships to form a comprehensive heterogeneous graph structure;

[0115] S33. Build a GAT embedding aggregation module: Use the GAT model to perform adaptive weight aggregation on the nodes in the heterogeneous graph to generate a user embedding representation that combines various information;

[0116] Through the above processing, each attribute of the user has been converted into a feature vector, including the user's portrait embedding e p 、the user's tweet list embedding e c 、the user's style embedding e style and the user's interest embedding e interestDifferent embedding features have different weights. Therefore, when the model aggregates these features, it needs to assign different weights to different types of features. Most of the existing models assign an average weight to each feature, or assign different weights to different types of features based on personal subjective feelings, or assign weights based on some statistics or experience. These methods all have a certain degree of subjectivity, and for the scenario of heterogeneous social network graphs with many types of features, this manual weight assignment method is very cumbersome and not conducive to the addition of new features.

[0117] To be able to assign an objective weight to different features and at the same time avoid manually assigning weights to features, the model uses the GAT model to aggregate these heterogeneous types of embedding features. The GAT model can assign weights to the embedding features according to the importance of the embedding features when aggregating the feature embeddings. Therefore, the user embedding features we finally obtain are an accurate embedding with adaptive weights, which can accurately describe the user. To be able to use the GAT model to aggregate the embedding features of users, first the model first embeds the user's portrait e p , the user's tweet list embedding e c , the user's style embedding e style and the user's interest embedding e interest into a heterogeneous graph, and establish edges for the relationships between the embedding nodes. Here, the model uses the user's portrait embedding node e p as the central node to connect the other three nodes e c , e style and e interest , thus completing the transformation of the embedding feature nodes into a heterogeneous graph. As shown in Equation 4, this paper uses g to represent the heterogeneous social network graph centered on the user, which includes the user's embedding feature nodes e p , e c and e style and e interest , as well as the edges e p → e c , e p → e style and e p → e interest .

[0118] g = ({e p , e c , e style e interest}, {e p → e c , e p → e style , e p → e interest})(Formula 4)

[0119] Then, the model feeds the heterogeneous graph g based on the user embedding features into the GAT network model to obtain a list of user embeddings that aggregates other embedding features. Here, the embeddings include the embeddings of each embedding feature node, and the model only selects the embedding of the central node e p as the final embedding representation of the user. As shown in Formula 5, e U represents the final user embedding representation, and Get represents obtaining the embedding representation of the final e p node obtained by GAT.

[0120] e U = Get(GAT(g), e p ) (Formula 5)

[0121] S34. Feature Fusion and Optimization Module: Fuse the features extracted by BERT, BiLSTM, and BERTopic, and finally generate a unified user embedding representation that combines various heterogeneous information of the user. This representation can comprehensively reflect the user's behavior and features in the social network.

[0122] Optionally, perform adaptive weight aggregation on the information nodes of each user heterogeneous type in the heterogeneous graph to obtain the model input vector, including:

[0123] Based on the graph attention model, set corresponding weight values for the nodes of the embedding vectors, content embedding vectors, and time information embedding vectors of each text attribute respectively;

[0124] Determine the model input vector according to the embedding vector and the corresponding weight value.

[0125] Specifically, the user attribute enhancement and unified embedding representation steps in S13 include:

[0126] Construct a text style and interest extraction module: Use the BERTopic model to perform topic modeling on the text content generated by the user, extract the topic information in the text, and convert it into an embedding vector;

[0127] Construct a heterogeneous graph module: Use various attribute embeddings, style embeddings, and interest embeddings of the user as nodes, and build edges between nodes through social relationships and content relationships to form a comprehensive heterogeneous graph structure;

[0128] Construct a GAT embedding aggregation module: Use the GAT model to perform adaptive weight aggregation on the nodes in the heterogeneous graph to generate a user embedding representation that combines multiple types of information;

[0129] Feature Fusion and Optimization Module: It fuses the features extracted by BERT, BiLSTM, and BERTopic, and finally generates a unified user embedding representation that synthesizes various heterogeneous information of the user. This representation can comprehensively reflect the user's behaviors and characteristics in the social network. The user identity alignment step in S15 includes:

[0130] After completing the unified user embedding representation, use the trained MLP model and input the unified user embedding representations in the two platforms as features;

[0131] The MLP model makes identity alignment predictions based on the learned user features to determine whether two users have the same identity.

[0132] The embodiment of this application also includes steps to optimize the model performance:

[0133] Adjust the model parameters according to the evaluation results, including but not limited to the number of layers of the GNN, the node embedding dimension, the dimension of the time embedding vector, the number of layers and the hidden layer size of the BERT model, and the selection of the optimizer and the dynamic adjustment strategy of the learning rate, etc., to improve the overall performance of the model in the user identity alignment task.

[0134] Evaluate and improve the aligned user identity results to ensure the accuracy and reliability of the alignment results. This module receives the aligned user identity data from the user identity alignment module and uses predefined evaluation metrics (including accuracy, F1 score, precision, recall, etc.) to objectively evaluate the alignment effect. These metrics can help evaluate the differences and similarities between the alignment results and the real data.

[0135] According to the evaluation results, the effect evaluation module can also propose improvement methods and suggestions, such as adjusting model parameters, retraining the model, or adopting other complementation strategies, etc., to further improve the quality and stability of the complementation results.

[0136] The embodiment of this application constructs a heterogeneous graph neural network model, comprehensively considers various heterogeneous information of users, and uses deep learning technology to encode and fuse this information to achieve more accurate and efficient user identity alignment. The embodiment of this application not only improves the accuracy of user identity alignment, but also enhances the adaptability and generalization ability of the model to data from different types of social platforms. Through the automated feature extraction and weight allocation mechanism, the embodiment of this application also significantly improves the efficiency of processing large-scale social network data, provides an effective technical means for social network analysis and user identity management, and can comprehensively process heterogeneous social information and effectively identify users' cross-platform identities, so as to more accurately simulate and predict user behaviors, and improve the personalization level of social network services and the accuracy of data mining.

[0137] The cross-platform user identity alignment processing method provided by the embodiments of the present application obtains the user behavior data of multiple users on different platforms; determines the identity information corresponding to the user behavior data according to a pre-trained identity alignment model, where the pre-trained identity alignment model is obtained by using the sample user data and the initial training model on different platforms to obtain the embedded vectors of heterogeneous information, encoding the embedded vectors of heterogeneous information to obtain the encoding result, and performing adaptive weight aggregation on the encoding result; determines the alignment information of multiple users according to the identity information, constructs a heterogeneous graph neural network model, comprehensively considers various heterogeneous information of users, and uses deep learning technology to encode and fuse this information to achieve more accurate and efficient user identity alignment, not only improving the accuracy of user identity alignment, but also enhancing the adaptability and generalization ability of the model to data on different types of social platforms. Through the automated feature extraction and weight assignment mechanism, the present invention also significantly improves the efficiency of processing large-scale social network data, providing an effective technical means for social network analysis and user identity management.

[0138] Another embodiment of the present application provides a cross-platform user identity alignment processing device for executing the cross-platform user identity alignment processing method provided by the above embodiment.

[0139] Referring to Figure 5 , a structural block diagram of an embodiment of a cross-platform user identity alignment processing device of the present application is shown. The device may specifically include the following modules: an acquisition module 501, a determination module 502, and an alignment module 503, where:

[0140] The acquisition module 501 is used to acquire the user behavior data of multiple users on different platforms;

[0141] The determination module 502 is used to determine the identity information corresponding to the user behavior data according to a pre-trained identity alignment model, where the pre-trained identity alignment model is obtained by using the sample user data and the initial training model on different platforms to obtain the embedded vectors of heterogeneous information, encoding the embedded vectors of heterogeneous information to obtain the encoding result, and performing adaptive weight aggregation on the encoding result;

[0142] The alignment module 503 is used to determine the alignment information of multiple users according to the identity information.

[0143] The cross-platform user identity alignment processing device provided by the embodiments of the present application obtains the user behavior data of multiple users on different platforms; determines the identity information corresponding to the user behavior data according to a pre-trained identity alignment model, where the pre-trained identity alignment model is obtained by using the sample user data and an initial training model on different platforms to obtain the embedded vectors of heterogeneous information, encoding the embedded vectors of heterogeneous information to obtain an encoding result, and performing adaptive weight aggregation on the encoding result; determines the alignment information of multiple users according to the identity information, constructs a heterogeneous graph neural network model, comprehensively considers various heterogeneous information of users, and uses deep learning technology to encode and fuse this information to achieve more accurate and efficient user identity alignment, not only improving the accuracy of user identity alignment, but also enhancing the adaptability and generalization ability of the model to data of different types of social platforms. Through the automated feature extraction and weight assignment mechanism, the present invention also significantly improves the efficiency of processing large-scale social network data, providing an effective technical means for social network analysis and user identity management.

[0144] Another embodiment of the present application further supplements the cross-platform user identity alignment processing device provided in the above embodiment.

[0145] Optionally, the device further includes a training module, and the training module is used for:

[0146] Obtain sample user data on different platforms;

[0147] Determine a user heterogeneous information dataset according to the sample user data;

[0148] Encode the heterogeneous information in the user heterogeneous information dataset based on a deep learning network model to obtain embedded vectors of heterogeneous information, where the embedded vectors of heterogeneous information at least include embedded vectors of text attributes, content embedded vectors, and time information embedded vectors;

[0149] Construct a heterogeneous graph according to the embedded vectors of heterogeneous information;

[0150] Perform adaptive weight aggregation on the information nodes of each user heterogeneous type in the heterogeneous graph to obtain a model input vector;

[0151] Input the model input vectors of users on different platforms into a multi-layer perceptron model respectively to obtain output results corresponding to the model input vectors of users on different platforms;

[0152] Determine whether the sample user data on different platforms comes from the same user according to the output results corresponding to the model input vectors of users on different platforms.

[0153] Optionally, the training module is used for:

[0154] The BERT model is used to encode the basic attribute information in the sample user data to obtain the embedding vectors of text attributes; among them, the basic attribute information includes at least the user name and nickname;

[0155] The BERTopic model is used to encode the user-generated content in the sample user data to obtain content embedding vectors, where the content embedding vectors include at least the topic information embedding vectors and the interest information embedding vectors; the timestamp information in the sample user data is encoded to obtain the time information embedding vectors;

[0156] The embedding vectors of text attributes, content embedding vectors, and time information embedding vectors are fused to obtain the embedding vectors of multi-dimensional heterogeneous information.

[0157] Optionally, the training module is used for:

[0158] Taking the embedding vectors of text attributes, content embedding vectors, and time information embedding vectors as nodes, and taking the relationship between the social relationship and the content relationship as edges, to determine a heterogeneous graph.

[0159] Optionally, the training module is used for:

[0160] Based on the graph attention model, corresponding weight values are set for the nodes of the embedding vectors of each text attribute, content embedding vectors, and time information embedding vectors;

[0161] According to the embedding vectors and the corresponding weight values, the model input vectors are determined.

[0162] For the apparatus embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, please refer to the partial description of the method embodiment.

[0163] The processing device for cross-platform user identity alignment provided by the embodiments of the present application obtains the user behavior data of multiple users on different platforms; determines the identity information corresponding to the user behavior data according to a pre-trained identity alignment model, where the pre-trained identity alignment model is obtained by using the sample user data and the initial training model on different platforms to obtain the embedded vectors of heterogeneous information, encoding the embedded vectors of heterogeneous information to obtain the encoding result, and performing adaptive weight aggregation on the encoding result; determines the alignment information of multiple users according to the identity information, constructs a heterogeneous graph neural network model, comprehensively considers various heterogeneous information of users, and uses deep learning technology to encode and fuse this information to achieve more accurate and efficient user identity alignment, not only improving the accuracy of user identity alignment, but also enhancing the adaptability and generalization ability of the model to data of different types of social platforms. Through the automated feature extraction and weight allocation mechanism, the present invention also significantly improves the efficiency of processing large-scale social network data, providing an effective technical means for social network analysis and user identity management.

[0164] Another embodiment of the present application provides a terminal device for executing the cross-platform user identity alignment processing method provided by the above embodiment.

[0165] Figure 6 is a schematic structural diagram of a terminal device of the present application, as Figure 6 shown, the terminal device includes: at least one processor 601 and a memory 602;

[0166] The memory stores a computer program; at least one processor executes the computer program stored in the memory to implement the cross-platform user identity alignment processing method provided by the above embodiment.

[0167] The terminal device provided in this embodiment obtains the user behavior data of multiple users on different platforms; determines the identity information corresponding to the user behavior data according to a pre-trained identity alignment model, where the pre-trained identity alignment model is obtained by using the sample user data and the initial training model on different platforms to obtain the embedded vectors of heterogeneous information, encoding the embedded vectors of heterogeneous information to obtain an encoding result, and performing adaptive weight aggregation on the encoding result; determines the alignment information of multiple users according to the identity information, constructs a heterogeneous graph neural network model, comprehensively considers various heterogeneous information of users, and uses deep learning technology to encode and fuse this information to achieve more accurate and efficient user identity alignment, which not only improves the accuracy of user identity alignment, but also enhances the adaptability and generalization ability of the model to data on different types of social platforms. Through the automated feature extraction and weight assignment mechanism, the present invention also significantly improves the efficiency of processing large-scale social network data, providing an effective technical means for social network analysis and user identity management.

[0168] Another embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed, it implements the cross-platform user identity alignment processing method provided in any of the above embodiments.

[0169] According to the computer-readable storage medium of this embodiment, the user behavior data of multiple users on different platforms is obtained; the identity information corresponding to the user behavior data is determined according to a pre-trained identity alignment model, where the pre-trained identity alignment model is obtained by using the sample user data and the initial training model on different platforms to obtain the embedded vectors of heterogeneous information, encoding the embedded vectors of heterogeneous information to obtain an encoding result, and performing adaptive weight aggregation on the encoding result; the alignment information of multiple users is determined according to the identity information, constructs a heterogeneous graph neural network model, comprehensively considers various heterogeneous information of users, and uses deep learning technology to encode and fuse this information to achieve more accurate and efficient user identity alignment, which not only improves the accuracy of user identity alignment, but also enhances the adaptability and generalization ability of the model to data on different types of social platforms. Through the automated feature extraction and weight assignment mechanism, the present invention also significantly improves the efficiency of processing large-scale social network data, providing an effective technical means for social network analysis and user identity management.

[0170] It should be noted that the above detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0171] Note that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly dictates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of the stated features, steps, operations, devices, components, and / or combinations thereof.

[0172] It should be noted that the terms "first", "second", etc. in the description and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein.

[0173] In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0174] For ease of description, spatial relative terms such as "above", "over", "on the upper surface", "above", etc. may be used herein to describe the spatial positional relationship of one device or feature to other devices or features as shown in the figures. It should be understood that the spatial relative terms are intended to encompass different orientations in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figure is inverted, the device described as "above" or "over" other devices or structures will then be positioned "below" or "beneath" other devices or structures. Thus, the exemplary term "above" can include both the orientation of "above" and "below". The device may also be positioned in other different ways, such as rotated 90 degrees or in other orientations, and corresponding interpretations of the spatial relative descriptions used herein will be made.

[0175] In the above detailed description, reference has been made to the accompanying drawings, which form a part hereof. In the drawings, like reference numerals typically identify like components, unless the context indicates otherwise. The illustrated embodiments described in the detailed description, the drawings, and the claims are not meant to be limiting. Other embodiments may be used and other changes may be made without departing from the spirit or scope of the subject matter presented herein.

[0176] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and variations can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A cross-platform processing method for user identity alignment, characterized in that The method includes: Obtaining user behavior data of multiple users on different platforms; Determining identity information corresponding to the user behavior data according to a pre-trained identity alignment model; wherein, the pre-trained identity alignment model is obtained by using sample user data and an initial training model on different platforms to obtain embedded vectors of heterogeneous information, encoding the embedded vectors of the heterogeneous information to obtain an encoding result, and performing adaptive weight aggregation on the encoding result; Determining alignment information of the multiple users according to the identity information; The identity alignment model is obtained by the following method: Obtaining sample user data on different platforms; Determining a user heterogeneous information dataset according to the sample user data; Encoding heterogeneous information in the user heterogeneous information dataset based on a deep learning network model to obtain embedded vectors of heterogeneous information, wherein the embedded vectors of heterogeneous information at least include embedded vectors of text attributes, content embedded vectors, and time information embedded vectors; Constructing a heterogeneous graph according to the embedded vectors of heterogeneous information; Performing adaptive weight aggregation on information nodes of each user heterogeneous type in the heterogeneous graph to obtain a model input vector; Inputting the model input vectors of users on different platforms into a multi-layer perceptron model respectively to obtain output results corresponding to the model input vectors of users on different platforms; Determining whether the sample user data on different platforms comes from the same user according to the output results corresponding to the model input vectors of users on different platforms.

2. The cross-platform user identity alignment processing method according to claim 1, wherein The encoding of heterogeneous information in the user heterogeneous information dataset based on a deep learning network model to obtain embedded vectors of heterogeneous information includes: Encoding basic attribute information in the sample user data by using a BERT model to obtain embedded vectors of text attributes; wherein, the basic attribute information at least includes a username and a nickname; Encoding user-generated content in the sample user data by using a BERTopic model to obtain content embedded vectors, wherein the content embedded vectors at least include embedded vectors of topic information and interest information; encoding timestamp information in the sample user data to obtain time information embedded vectors; Performing fusion processing on the embedded vectors of text attributes, the content embedded vectors, and the time information embedded vectors to obtain multi-dimensional embedded vectors of heterogeneous information.

3. The cross-platform user identity alignment processing method according to claim 1, characterized in that The constructing a heterogeneous graph according to the embedded vectors of heterogeneous information includes: Using the embedded vectors of text attributes, the content embedded vectors, and the time information embedded vectors as nodes, and using the relationship between social relationships and content relationships as edges to determine the heterogeneous graph.

4. The processing method for cross-platform user identity alignment according to claim 1, characterized in that The performing adaptive weight aggregation on information nodes of each user heterogeneous type in the heterogeneous graph to obtain a model input vector includes: Based on a graph attention model, setting corresponding weight values for nodes of each embedded vector of text attributes, the content embedded vector, and the time information embedded vector respectively; Determining the model input vector according to the embedded vectors and the corresponding weight values.

5. A cross-platform processing device for user identity alignment, characterized in that, The device includes: An acquisition module, configured to acquire user behavior data of multiple users on different platforms; A determination module, configured to determine identity information corresponding to the user behavior data according to a pre-trained identity alignment model; wherein, the pre-trained identity alignment model is obtained by using sample user data on different platforms and an initial training model to obtain embedding vectors of heterogeneous information, encoding the embedding vectors of the heterogeneous information to obtain an encoding result, and performing adaptive weight aggregation on the encoding result; An alignment module, configured to determine alignment information of the multiple users according to the identity information; The apparatus further includes a training module, and the training module is configured to: Acquire sample user data on different platforms; Determine a user heterogeneous information dataset according to the sample user data; Encode heterogeneous information in the user heterogeneous information dataset based on a deep learning network model to obtain embedding vectors of the heterogeneous information, wherein the embedding vectors of the heterogeneous information at least include embedding vectors of text attributes, content embedding vectors, and time information embedding vectors; Construct a heterogeneous graph according to the embedding vectors of the heterogeneous information; Perform adaptive weight aggregation on information node information of each user heterogeneous type in the heterogeneous graph to obtain a model input vector; Respectively input the model input vectors of users on different platforms into a multi-layer perceptron model to obtain output results corresponding to the model input vectors of users on different platforms; Determine whether the sample user data on different platforms comes from the same user according to the output results corresponding to the model input vectors of users on different platforms.

6. The cross-platform user identity alignment processing device according to claim 5, characterized in that The training module is configured to: Use a BERT model to encode basic attribute information in the sample user data to obtain embedding vectors of text attributes; wherein, the basic attribute information at least includes a user name and a nickname; Use a BERTopic model to encode user-generated content in the sample user data to obtain content embedding vectors, wherein the content embedding vectors at least include topic information embedding vectors and interest information embedding vectors; encode timestamp information in the sample user data to obtain time information embedding vectors; Perform fusion processing on the embedding vectors of the text attributes, the content embedding vectors, and the time information embedding vectors to obtain multi-dimensional embedding vectors of heterogeneous information.

7. A terminal device, characterized in that, Including: At least one processor and a memory; The memory stores a computer program; The at least one processor executes the computer program stored in the memory to implement the cross-platform user identity alignment processing method according to any one of claims 1-4.

8. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed, the cross-platform user identity alignment processing method according to any one of claims 1-4 is implemented.

Citation Information

Patent Citations

  • Modeling method of network community class group and user representation model

    CN117312489A