A Cross-Social Network User Alignment Method Based on Hybrid Extended Cognitive Features
The hybrid cognitive feature expansion method using large language models and graph attention networks enhances cross-platform user alignment by leveraging deep semantic understanding and user profiling to address data variability and heterogeneity, improving alignment accuracy.
Patent Information
- Application Number
- CN202411956952.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-29
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-12-29
AI Technical Summary
There are problems in the user alignment task of cross-social networks, data privacy, data heterogeneity and data sparsity, resulting in low user alignment efficiency and accuracy.
Using a method based on hybrid extension cognitive characteristics, the original data of the user account is expanded using a large language model, the fine-tuned pre-trained language model DeBERTaV3 is encoded, and the user node graph is constructed using the graph neural network model GAT, the node embedding vector is calculated, and the user node pair is finally matched through the distance matrix.
It improves the efficiency and accuracy of user alignment across social networks, solves problems such as variable data structures, strong content heterogeneity, and sparse network structures, and achieves more accurate user alignment.
Smart Images

Figure CN119939265B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of artificial intelligence and social network data mining, and particularly relates to a cross-social network user alignment method based on hybrid extended cognitive features. Background Art
[0002] Due to the significant differences in functions and content characteristics among different social platforms, the same natural person user often has accounts on multiple different social platforms to meet different needs, and will post different content on different social platforms. Users who register accounts in different networks can act as bridges connecting different networks, thereby connecting and integrating multiple social networks to achieve the following purposes: (1) Construct a comprehensive user feature representation: By integrating user information on multiple social network platforms, a more comprehensive and accurate user feature representation is constructed, which helps to better understand user needs and behaviors. (2) Achieve cross-network recommendation: Based on the behaviors and interests of users on different platforms, cross-network friend recommendation, content recommendation, etc. are realized to improve the user experience. The cross-social network user alignment task refers to the process of finding accounts belonging to the same natural person account among many accounts on different social network platforms.
[0003] The main difficulties in realizing the cross-social network user alignment task are as follows: (1) Data privacy issue: Due to privacy protection and security issues, some unique user attributes (such as email addresses, mobile phone numbers, etc.) are difficult to obtain, and users cannot be directly identified through these unique attributes. (2) Data heterogeneity issue: User data on different social network platforms often have huge differences. For example, the content on Twitter is instant and public, and is limited to short texts (within 280 characters); while the content on Facebook is mainly social interaction, and longer texts (within 63206 characters) can be published. This poses a great challenge to user alignment. (3) Data sparsity issue: The user scale of real social networks is huge, while the accounts that can be marked as aligned users are very few in comparison. The sparsity of sample data adds great difficulty to data-based solutions. Summary of the Invention
[0004] (1) Technical Problems to be Solved
[0005] The technical problem to be solved by the present invention is to provide a cross-social network user alignment method to improve the efficiency and accuracy of the cross-social network user alignment task.
[0006] (2) Technical Solutions
[0007] To solve the above technical problems, the present invention provides a cross-social network user alignment method based on hybrid extended cognitive features, including the following steps:
[0008] (1) Expansion of the original data of user accounts based on large language models: Design the prompt words, and use the designed prompt words to apply the GLM-4-9B large language model to analyze the original data of user accounts, and use the output of the large language model as the expanded user feature data;
[0009] (2) Encoding using a fine-tuned pre-trained language model: Fine-tune the pre-trained language model DeBERTaV3, and input the original data and the expanded user feature data in step (1) into the fine-tuned pre-trained language model DeBERTa V3 for encoding to obtain the embedded representation vector in text form for each user;
[0010] (3) Obtaining the embedded representation of the nodes of the graph using a graph neural network model: Train the graph neural network model. The graph neural network model selects the graph attention model GAT. Use the embedded representation vector of each user obtained in step (2) as the feature vector of the node, and use the mutual relationship between user accounts as the mutual relationship between nodes to construct a user node graph, and input the user node graph into the trained graph neural network model to obtain the node embedding vector of each node;
[0011] (4) Matching the graph node pairs: Use the node embedding vectors obtained after processing the original data of user accounts on different social networks through steps (1), (2), and (3) to construct a distance matrix, and infer the matching aligned user node pairs according to the distance matrix.
[0012] The present invention also provides a system for implementing the above method.
[0013] (III) Advantageous effects
[0014] The present invention proposes a cross-social network user alignment method based on hybrid extended cognitive features, which uses a large language model to expand three types of cognitive psychological features such as personality features, language features, and ability features of the original data of social network users, and realizes more accurate cross-social network user alignment based on the hybrid features of the original attribute features and cognitive psychological features. Using hybrid extended cognitive features solves to a certain extent the problems of variable data structures, strong heterogeneity of published content, and sparse network structures in the cross-social network user alignment task. Description of the drawings
[0015] Figure 1 is the schematic diagram of the method of the present invention;
[0016] Figure 2 is the schematic diagram of analyzing the original data of user accounts in the method of the present invention. Detailed implementation manners
[0017] To make the objectives, content, and advantages of the present invention clearer, the following further describes in detail the specific implementation manners of the present invention in conjunction with the accompanying drawings and embodiments.
[0018] The present invention provides a cross-social network user alignment method based on hybrid extended cognitive features. As an emerging technology that has developed rapidly in recent years, the success of large language models in many fields has proven their powerful context understanding ability and semantic reasoning ability. Applying large language models to perform in-depth semantic understanding and information mining of social network content has great potential. Therefore, aiming at the difficulties in the above-mentioned cross-social network user alignment task, the present invention uses large language models to mine the deep feature information of social network accounts to improve the efficiency and accuracy of the cross-social network user alignment task.
[0019] This method mainly includes the following four steps: Step 1, use a large language model to process the original data of different social network user accounts to obtain extended cognitive features; Step 2, use a fine-tuned pre-trained language model to encode and obtain the embedded representation vectors of the original attribute features and extended cognitive features of the user accounts; Step 3, use the above embedded representation vectors as node information, and establish the mutual relationship of nodes based on the mutual relationship (such as mutual following, friend relationship, etc.) of different user accounts on the social network, construct a social network graph representation, input different social network graphs into the graph neural network model, and obtain the node embedded representation of the graph; Step 4, calculate the distance matrix of the node embedded representations of two different social network graphs, and predict the node pairs of the user alignment task.
[0020] Specifically, referring to Figure 1 , the cross-social network user alignment method based on data augmentation using large language models proposed by the present invention includes the following steps:
[0021] (1) Expansion of the original data of user accounts based on large language models: Analyze the original data of user accounts by applying the GLM-4-9B large language model through designed prompts, and use the output of the large language model as the extended user feature data;
[0022] (2) Encoding using a fine-tuned pre-trained language model: Fine-tune the pre-trained language model DeBERTaV3, and input the original data and the extended user feature data in step (1) into the fine-tuned pre-trained language model DeBERTa V3 for encoding to obtain the embedded representation vectors in text form (i.e., embedded vectors) of each user;
[0023] (3) Obtain the embedded representations of the nodes of the graph using a graph neural network model: Train the graph neural network model. The graph neural network model selects the Graph Attention Network (GAT). Use the embedded representation vectors of each user obtained in step (2) as the feature vectors (node information) of the nodes, and use the mutual relationships between user accounts (such as following, friendship, etc.) as the mutual relationships between nodes to construct a user node graph (as a social network graph). Input the user node graph into the trained graph neural network model to obtain the node embedded vectors of each node;
[0024] (4) Match the graph node pairs: Use the node embedded vectors obtained after processing the original data of user accounts on different social networks through steps (1), (2), and (3) to construct a distance matrix, and infer the matching aligned user node pairs based on the distance matrix.
[0025] In step (1), use the designed prompt words to enhance the original data of user accounts using a large language model. The main purpose is to expand the limited user information and find the invariant deep information behind the variable surface data of different social network accounts of the same natural person. The original data of user accounts includes the following content: <1> Account attributes of the user: username, gender, birthday, industry, educational experience, personal profile; due to the design of different platforms and the filling preferences of users, account attributes are often missing; <2> Content published by the user: text data of the user's historical tweets. The output of the large language model refers to the answer given by the large language model after analyzing the original data of the user account according to the requirements put forward by the prompt words, including a data type obtained through discrimination and the basis for discrimination, both presented in text form.
[0026] The content posted by the same user on different platforms often varies greatly, which makes it difficult to directly compare the original texts posted by the user for user alignment. However, these variable contents largely depend on the stable cognitive characteristics of a natural person. In this step, a large language model is applied in combination with cognitive psychology theory, and relatively mature personality theory, ability theory, and language theory are selected to analyze the original data of user accounts. Among them, the Myers-Briggs Type Indicator (MBTI) is selected for analysis in the personality theory. The Myers-Briggs Type Indicator is a personality type theory model that divides personality into four groups of opposite innate preferences (introversion and extraversion, sensing and intuition, thinking and feeling, judging and perceiving, and the four preferences can form 16 stable personality types). The ability theory selects the occupation, an explicit representation, for analysis. The occupation classification is based on the Standard Occupational Classification (SOC) system in the United States, which divides 1,016 social occupations into four levels: major group, minor group, broad occupation, and SOC title, and can clearly distinguish the characteristics of different occupations, such as job content, experience requirements, skill types, and professional fields. The language theory analyzes the dialects and variant grammars of social network users to conduct a representational analysis of deep cognitive characteristics.
[0027] Reference Figure 2 , in the personality theory of cognitive psychology, a person's personality will affect the expression style of their tweets on social platforms. For example, personality types with a stronger rational tendency (such as INTP, ENTJ, etc.) may be more inclined to express their views directly and objectively, and use fewer emotionally charged words. Personality types with a stronger emotional tendency (such as ENFP, INFP, etc.) may be better at using emotional language to express their feelings and views, making the content of their speeches more contagious. Extroverted personalities (such as ESTP, ENFP) may be better at incorporating humorous elements into their speeches, making the communication atmosphere more relaxed and pleasant. Introverted or judging personalities (such as INTJ, ISTJ) may pay more attention to the accuracy and depth of the content, and their speech styles are relatively more serious; personality also affects their emotional expression on social platforms. For example, extroverted or emotional personalities (such as ENFP, ESFJ) may be more willing to share their emotional experiences and inner worlds on social platforms to seek resonance with others. Introverted or thinking personalities (such as INTP, INFJ) may be relatively more reserved and less likely to express personal emotions in public.
[0028] In cognitive psychology's language theory, the dialects and variants of English around the world (such as British English, American English, Australian English, New Zealand English, Singaporean English, Indian English, etc.) often have their own unique vocabulary and grammar. For example, in Indian English, there is often a mixture of Hindi and English (Hinglish), such as "I have hazaarthingson my mind right now" (I have a thousand things on my mind right now), where "hazaar" is a Hindi word meaning "a thousand", and the whole sentence uses the sentence structure of English; in Australian English, there are many unique local words, which reflect the geographical, climatic, flora and fauna, and cultural characteristics of Australia. "Thongs" means flip-flops in Australia, "barbie" means barbecue grill, and "macca’s" is a nickname for McDonald's. Some English words have changed in meaning or usage in Australian English. For example, "barbecue" in Australia not only refers to the barbecue activity, but is also often used to refer to the barbecue grill; "g'day" is a way of greeting, meaning "hello". By identifying these unique language features, the language variant a person habitually uses can be directly reflected in the content they post on social platforms.
[0029] In cognitive psychology's theory of ability, occupation is a prominent representation of ability. The occupation a person is engaged in affects the professionalism and relevance of the content they post. People in different occupations have different professional knowledge and skills in their fields, and these knowledge and skills will be reflected in their social platform speeches. For example, doctors may share health knowledge on the platform, while lawyers may discuss legal cases and viewpoints. This professionalism makes their speeches more authoritative and credible. The professional background also determines which topics an individual is more likely to focus on and discuss on social platforms. For example, technology workers may be more concerned about the development of new technologies and new products, while educators may be more concerned about education policies, teaching methods, etc.; the professional background also affects the stance and perspective of the content they post. For example, environmental protection workers may be more inclined to express their views on a certain policy or event from an environmental protection perspective, while entrepreneurs may be more concerned about the impact of the policy or event on the business environment.
[0030] Therefore, in this step, a large language model is used to deeply analyze the original data of social network user accounts, and the original data and user attributes are extended in three dimensions: personality characteristics, ability characteristics, and language characteristics to generate account mixed characteristics. The specific approach is to design prompts to query the large language model, and the specific design is as follows:
[0031] <1>Analysis of personality characteristics:
[0032] In terms of personality traits, the Myers-Briggs Type Indicator (MBTI) is used for analysis, and a large language model is utilized to analyze the user's personality type based on the text content they post. The following prompt is designed:
[0033] Here are some posts from a user posting on the internet. Please infer his / her MBTI type based on these posts. Your response should clearly provide one MBTI label and give reasons for each dimension in MBTI, posts:’......’
[0034] This prompt specifies analyzing a specific MBTI personality type and clearly requires providing specific analyses for the four preference dimensions.
[0035] <2>Analysis of language features:
[0036] In terms of language, a large language model is used to analyze the English varieties and dialects used in the text content posted by the user to uncover the language features (including special words, phrases, and grammar) behind the text information. The following prompt is designed:
[0037] Here are some posts from a user posting on the internet. Please infer what variety of the English language he / she uses. Your response should clearly provide one variety label and give reasons. If you can't make a clear inference, please give the answer "not sure", posts:’......’
[0038] This prompt specifies analyzing the English variety to which the language of a user's post belongs, gives the requirements for the analysis, and requires giving the answer "not sure" if there are no clear features, because the text posted by the user does not always carry the characteristics of a clear variant language.
[0039] <3>Analysis of ability traits - occupation:
[0040] In terms of ability characteristics, a large language model is used to analyze the user's occupation category based on the text content posted by the user; the occupation category is based on the Standard Occupational Classification (SOC) system in the United States. The design prompt is as follows:
[0041] Here are some posts from a user posting on the internet. Please infer his / her careers. Your response should clearly provide one career title in Standard Occupational Classification (SOC) and give reasons. If you can't make a clear inference, please give the answer "not sure", posts: ’......’
[0042] This prompt specifies the requirement to analyze the possible occupation category and give reasons from the text content posted on a user's social network, and requires giving the answer "not sure" if there are no clear characteristics, because the user's occupation is not always clearly reflected in the text posted.
[0043] The fine-tuning of the pre-trained language model DeBERTaV3 in step (2) and the training of the graph attention model GAT in step (3) are implemented using the same training method, that is, using the contrastive learning method to make the samples of the same user closer in the encoding space, while the samples of different users are farther away. The specific training method is as follows:
[0044] <1> Forward propagation: Input the data into the model to obtain the output of the model. The pre-trained language model DeBERTaV3 obtains the word embedding vectors of each user account, and the graph attention model GAT obtains the (graph) node embedding vectors of each user account;
[0045] <2> Calculate the loss function: The info-NCE loss function is used as the loss function for fine-tuning the model using the contrastive learning method. The specific formula is as follows:
[0046]
[0047]
[0048] where, l iis the loss value of a single sample, L is the average loss value of a mini-batch of samples, M is the total number of samples in the mini-batch, and z i and are respectively the i-th user account sample in a social network and the account sample belonging to the same user in another social network corresponding to it, and P pos is the set of user accounts, and τ is an adjustable hyperparameter;
[0049] <3>Backpropagation: Calculate the gradient of the loss value with respect to the model parameters according to the info-NCE loss function, and update the model parameters according to the loss value using the Adam optimization algorithm;
[0050] <4>Repeat iteration: Repeatedly execute the above processes of forward propagation, calculating the loss function, and backpropagation until the model reaches the specified number of iterations.
[0051] In step (4), the pairwise distances are calculated for the node embedding vectors (node feature vectors) of all users in the two social networks after being processed by steps (1), (2), and (3). The pairwise distances between the node embedding vectors are calculated using the cosine distance, and the specific formula is as follows:
[0052]
[0053] where u1 and u2 are the feature vectors of two user nodes respectively.
[0054] After calculating the pairwise distances for all pairs of node embedding vectors of user nodes in the two social networks, a distance matrix D of size N×M is obtained, where N and M are the numbers of user nodes in the two social networks respectively. For each user node to be predicted, the pairwise distances of the node pairs formed by it and all nodes in the other social network are sorted, and the node (the node in the other social network) with the smallest pairwise distance is used as the predicted aligned node. The user node to be predicted and the corresponding aligned node form an aligned user node pair, which is used as the aligned account result.
[0055] It can be seen that the technical solution of the present invention is divided into steps such as cognitive feature expansion, encoding using a pre-trained language model, obtaining node embedding representations using a graph neural network, and matching graph node pairs, and involves the following key technologies:
[0056] 1) Cognitive feature expansion
[0057] Due to the differences in the functions, community atmospheres, and habits of different social platforms, the content shown by the accounts of the same natural user on different social platforms often varies greatly. However, the cognitive characteristics such as the personality characteristics, language characteristics, and ability characteristics of the user behind the account are often fixed. Therefore, the context understanding ability and semantic reasoning ability of the large language model can be used to analyze the deep cognitive characteristics of social account users, and extended characteristics of the original account data can be obtained.
[0058] 2) Use a pre-trained language model for encoding
[0059] For the original data and extended data of social network accounts represented in text form, use the pre-trained language model of RoBERTaV3 to convert the text into fixed-size embedding vectors to encode the semantic information in the text into the vector space for subsequent task use.
[0060] 3) Use a graph neural network to obtain node embedding representations
[0061] In the task of cross-social network user alignment, not only the characteristics of the user nodes themselves, but also the characteristics of the users associated with the user nodes can provide features for analysis. Therefore, use a graph neural network model to encode the node features associated with each node into the node's embedding representation using its message passing mechanism between nodes.
[0062] 4) Match the graph node pairs
[0063] Calculate the distance between the embedding vectors of the nodes of the two graphs for each node, and after obtaining the distance matrix, sort the distance of the feature vectors of the matching nodes of each node to be predicted in the other network to predict the aligned user nodes.
[0064] The following embodiments specifically elaborate on the cross-social network user alignment method based on the hybrid extended cognitive characteristics of the present invention.
[0065] The data used in this embodiment: Twenty thousand user account data that can be linked based on the mutual following relationship collected from social platform A and social platform B respectively, among which 100 user data can be aligned with the other platform. The data examples are as follows in the table:
[0066]
[0067]
[0068] 1. Use a large language model for data augmentation (expansion): Ask the large model with the above-mentioned three types of prompt words for the user's tweet text, and use the answer content as a supplement to obtain an extended dataset. The extended data examples are as follows in the table:
[0069]
[0070] 2. Training of the model: Divide the data into a training set and a validation set in a ratio of 8:2, and train the DeBERTaV3 language model and the GAT graph neural network model in a contrastive learning manner. Select the model parameters with the best performance on the validation set during 5000 rounds of training as the final model parameters.
[0071] 3. Inference of the model: Input the 500 user data of each of the two social networks to be recognized into the large language model GLM-4-9B to obtain the extended data (data format: text); use the extended data as the input of the fine-tuned DeBERTaV3 language model, and the output is the word embedding vector (data format: floating-point number array); use the word embedding vector and the user relationship adjacency matrix (data format: a two-dimensional boolean array) as the input of the graph neural network, and the output is the graph node embedding vector of each user (data format: floating-point number array).
[0072] Alignment of user account recognition: Calculate the cosine distance pairwise for the graph node embedding vectors of each user obtained from the two networks to obtain a distance matrix of size 500×500 (data format: two-dimensional floating-point matrix). For the target user account to be recognized, select all the distance values corresponding to its corresponding row or column in the distance matrix, sort them, and the user node corresponding to the shortest distance value is the aligned user node of the target user account on the other network.
[0073] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.
Claims
1. A cross-social network user alignment method based on hybrid extended cognitive features, characterized in that Including the following steps: (1) Expansion of the original data of user accounts based on large language models: Design prompts, and use the designed prompts to apply the GLM-4-9B large language model to analyze the original data of user accounts. Take the output of the large language model as the expanded user feature data; (2) Encoding using a fine-tuned pre-trained language model: Fine-tune the pre-trained language model DeBERTaV3, and input the original data and the expanded user feature data in step (1) into the fine-tuned pre-trained language model DeBERTaV3 for encoding to obtain the embedded representation vector in text form for each user; (3) Obtaining the embedded representation of the nodes of the graph using a graph neural network model: Train the graph neural network model. The graph neural network model selects the graph attention model GAT. Use the embedded representation vector of each user obtained in step (2) as the feature vector of the node, and use the mutual relationship between user accounts as the mutual relationship between nodes to construct a user node graph. Input the user node graph into the trained graph neural network model to obtain the node embedding vector of each node; (4) Matching the graph node pairs: Use the node embedding vectors obtained after processing the original data of user accounts on different social networks through steps (1), (2), and (3) to construct a distance matrix, and infer the matching aligned user node pairs based on the distance matrix.
2. The method according to claim 1, wherein In step (1), using the designed prompts to apply the GLM-4-9B large language model to analyze the original data of user accounts can enhance the original data of user accounts, expand limited user information, and find the invariant deep information behind the variable surface data of different social network accounts of the same natural person. The original data of user accounts includes the following content: <1> Account attributes of the user: username, gender, birthday, industry, education experience, and personal profile; <2> Content published by the user: text data of the user's historical tweets. The output of the large language model refers to the answer given by the large language model after analyzing the original data of the user account according to the requirements put forward by the prompt, including a discriminated data type and the basis for discrimination, both in text form.
3. The method according to claim 1, characterized in that, In step (1), the large language model is applied in combination with cognitive psychology theory, and the personality theory, ability theory, and language theory are selected to analyze the original data of user accounts; among them, the Myers-Briggs Type Indicator (MBTI) is selected for analysis in the personality theory; the dominant representation of occupation is selected for analysis in the ability theory; the deep cognitive characteristics are represented and analyzed in the language theory by analyzing the dialects and variant grammars of social network users.
4. The method according to claim 3, wherein When applying the large language model in combination with cognitive psychology theory and selecting the personality theory, ability theory, and language theory to analyze the original data of user accounts in step (1), the original data and user attributes are expanded in three dimensions of personality characteristics, ability characteristics, and language characteristics to generate account mixed characteristics; specifically, design prompts to ask the large language model, and the specific design is as follows: <1>Analysis of personality traits: In terms of personality traits, the Myers-Briggs Type Indicator (MBTI) is used for analysis. The large language model is utilized to analyze the personality type of users based on the text content they post. The designed prompt words specify that a specific MBTI personality type needs to be analyzed, and reasons for the specific analysis of the four preference dimensions should be given. <2>Analysis of language features: In terms of language, the large language model is used to analyze the English variants and dialects used in the text content posted by users, in order to uncover the language features behind the text information. The designed prompt words require analyzing the English variant to which the language of a user's posted content belongs, and reasons for the analysis should be given. Also, if there are no clear features, an uncertain answer should be provided. <3>Analysis of ability traits: In terms of ability traits, the large language model is used to analyze the professional category of users based on the text content they post. The designed prompt words require analyzing the professional classification of a user from the text content posted on their social network and giving reasons. Also, if there are no clear features, an uncertain answer should be provided.
5. The method according to claim 1, wherein The fine-tuning of the pre-trained language model DeBERTa V3 in step (2) and the training of the graph attention model GAT in step (3) are implemented using the same training method, that is, using the contrastive learning method to make the samples of the same user closer in the encoding space, while the samples of different users are farther away.
6. The method according to claim 5, characterized in that, The following training method is used to implement the fine-tuning of the pre-trained language model DeBERTa V3 in step (2) and the training of the graph attention model GAT in step (3): <1>Forward propagation: Input the data into the model to obtain the output of the model. The pre-trained language model DeBERTa V3 obtains the word embedding vectors of each user account, and the graph attention model GAT obtains the node embedding vectors of each user account. <2>Calculate the loss function: The info-NCE loss function is used as the loss function for fine-tuning the model using the contrastive learning method. The specific formula is as follows: where, l i is the loss value of a single sample, L is the average loss value of a mini-batch of samples, M is the total number of samples in the mini-batch, z i and are the i-th user account sample in a social network and the account sample belonging to the same user in another corresponding social network respectively, P pos is the set of user accounts for the corresponding user pair, and τ is an adjustable hyperparameter; <3>Backward propagation: Calculate the gradient of the loss value with respect to the model parameters according to the info-NCE loss function, and use the Adam optimization algorithm to update the model parameters according to the loss value. <4>Repeat iteration: Repeat the above processes of forward propagation, calculating the loss function, and backward propagation until the model reaches the specified number of iterations.
7. The method according to claim 1, wherein In step (4), the pairwise distances are calculated for the node embedding vectors of all users in the two social networks after being processed by steps (1), (2), and (3). The pairwise distances between the node embedding vectors are calculated using the cosine distance. The specific formula is as follows: where u1 and u2 are the feature vectors of two user nodes respectively; After pairwise calculating the pairwise distances of the node embedding vectors of all user nodes in two social networks, a distance matrix D of size N×M is obtained, where N and M are the numbers of user nodes in the two social networks respectively; for each user node to be predicted, the pairwise distances of the node pairs formed by it and all nodes in the other social network are sorted, and the node with the smallest pairwise distance is used as the predicted aligned node, and the user node to be predicted and the corresponding aligned node form an aligned user node pair.
8. The method according to any one of claims 1 to 7, characterized in that, This method is applied in the fields of artificial intelligence and social network data mining technology.
9. A system for implementing the method according to any one of claims 1 to 7.
10. The system according to claim 9, wherein This system is applied in the fields of artificial intelligence and social network data mining technology.
Citation Information
Patent Citations
Cross-social network user identity recognition method and system based on machine learning
CN109753602A
Cross-network user alignment method based on deep learning
CN110347932A