Cross-social network user alignment method based on mixed extended cognitive features

By adopting a hybrid extended cognitive feature method in the user alignment task of cross-social networks, and using large language model and graph neural network model, data privacy, heterogeneity and sparseness problems in user alignment tasks are solved, and alignment efficiency and accuracy are improved.

CN119939265AActive Publication Date: 2025-05-06BEIJING INST OF COMP TECH & APPL
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411956952.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-29
Publication Date
2025-05-06
Estimated Expiration
2044-12-29

AI Technical Summary

Technical Problem

Cross-social network user alignment tasks face difficulties such as data privacy, data heterogeneity and data sparsity, resulting in low efficiency and accuracy.

Method used

Using a method based on hybrid extension cognitive features, the original data of the user account is expanded through a large language model, combined with fine-tuned pre-trained language model and graph neural network model, the user's embedded representation vector is obtained, and the user alignment is matched through the distance matrix.

Benefits of technology

It improves the efficiency and accuracy of the alignment tasks of cross-social network users, and solves problems such as variable data structures, strong content heterogeneity, and sparse network structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939265A_ABST
    Figure CN119939265A_ABST
Patent Text Reader

Abstract

The invention relates to a cross-social network user alignment method based on mixed extended cognitive features, and belongs to the technical field of artificial intelligence and social network data mining. According to the method, three types of cognitive psychological characteristics such as personality characteristics, language characteristics and ability characteristics of the original data of the social network users are expanded by adopting a large language model, and more accurate cross-social network user alignment is realized based on the mixed characteristics of the original attribute characteristics and the cognitive psychological characteristics. The problems of changeable data structure, high published content heterogeneity, sparse network structure and the like in a cross-social network user alignment task are solved to a certain extent by adopting the mixed extended cognitive features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence and social network data mining, and specifically relates to a cross-social network user alignment method based on hybrid extended cognitive features. Background Art

[0003] Since different social platforms have significant differences in functions and content characteristics, the same natural person user often has multiple accounts on different social platforms to meet different needs, and will publish different content on different social platforms. These users who register accounts in different networks can act as bridges connecting different networks, thereby connecting and integrating multiple social networks to achieve the following goals: (1) Construct a comprehensive user feature representation: By integrating user information on multiple social network platforms, a more comprehensive and accurate user feature representation is constructed, which helps to better understand user needs and behaviors. (2) Achieve cross-network recommendations: Based on the user's behavior and interests on different platforms, cross-network friend recommendations, content recommendations, etc. are achieved to improve user experience. The cross-social network user alignment task refers to the process of finding accounts belonging to the same natural person among many accounts on different social network platforms.

[0004] The implementation of the cross-social network user alignment task has the following main difficulties: (1) Data privacy issues: Due to privacy protection and security issues, some unique user attributes (such as email addresses, mobile phone numbers, etc.) are difficult to obtain, and users cannot be directly identified through these unique attributes. (2) Data heterogeneity issues: User data on different social network platforms often differ greatly. For example, the content on Twitter is instant, public, and limited to short texts (within 280 characters); while the content on Facebook is mainly social interaction, and longer texts (within 63206 characters) can be posted. This poses a huge challenge to user alignment. (3) Data sparsity issues: The user scale of real social networks is huge, while the accounts that can be marked as aligned users are relatively rare. The sparsity of sample data greatly increases the difficulty of data-based solutions. Summary of the invention

[0005] 1. Technical issues to be resolved

[0006] The technical problem to be solved by the present invention is to provide a method for aligning users across social networks to improve the efficiency and accuracy of the task of aligning users across social networks.

[0007] (II) Technical solution

[0008] In order to solve the above technical problems, the present invention provides a method for aligning users across social networks based on hybrid extended cognitive features, comprising the following steps:

[0009] (1) Expansion of user account raw data based on large language model: Design prompt words, and use the designed prompt words to apply the GLM-4-9B large language model to analyze the raw data of user accounts, and use the output of the large language model as the expanded user feature data;

[0010] (2) Encoding using a fine-tuned pre-trained language model: Fine-tune the pre-trained language model DeBERTa V3, input the original data and expanded user feature data in step (1) into the fine-tuned pre-trained language model DeBERTa V3 for encoding, and obtain the text-based embedding representation vector of each user;

[0011] (3) Use the graph neural network model to obtain the embedded representation of the nodes of the graph: Train the graph neural network model. The graph neural network model uses the graph attention model GAT. The embedded representation vector of each user obtained in step (2) is used as the feature vector of the node, and the relationship between user accounts is used as the relationship between nodes. The user node graph is constructed and the user node graph is input into the trained graph neural network model to obtain the node embedding vector of each node.

[0012] (4) Matching graph node pairs: Use the node embedding vectors obtained after processing the original data of user accounts on different social networks in steps (1), (2), and (3) to construct a distance matrix, and infer the matching aligned user node pairs based on the distance matrix.

[0013] The present invention also provides a system for implementing the method.

[0014] (III) Beneficial effects

[0015] This paper proposes a cross-social network user alignment method based on hybrid extended cognitive features, which uses a large language model to expand the three types of cognitive psychological features of the original data of social network users, such as personality features, language features, and ability features, and achieves more accurate cross-social network user alignment based on the hybrid features of original attribute features and cognitive psychological features. The use of hybrid extended cognitive features can solve the problems of variable data structure, strong heterogeneity of published content, and sparse network structure in the cross-social network user alignment task to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is the principle diagram of the method of the present invention;

[0017] Figure 2This is a schematic diagram of the principle of analyzing the original data of a user account in the method of the present invention. DETAILED DESCRIPTION

[0018] In order to make the purpose, content and advantages of the present invention more clear, the specific implementation methods of the present invention are further described in detail below in conjunction with the drawings and examples.

[0019] The present invention provides a method for aligning users across social networks based on hybrid extended cognitive features. As an emerging technology that has developed rapidly in recent years, the success of large language models in many fields has proved its powerful contextual understanding and semantic reasoning capabilities. The application of large language models to perform deep semantic understanding and information mining on social network content has great potential. Therefore, in view of the difficulties in the above-mentioned cross-social network user alignment task, the present invention uses large language models to mine deep feature information of social network accounts to improve the efficiency and accuracy of the cross-social network user alignment task.

[0020] The method mainly includes the following four steps: Step 1, using a large language model to process the original data of different social network user accounts to obtain extended cognitive features; Step 2, using a fine-tuned pre-trained language model to encode the original attribute features of the user account and the embedded representation vector of the extended cognitive features; Step 3, using the above embedded representation vector as node information, and using the mutual relationships between different user accounts on the social network (mutual attention, friend relationships, etc.) to establish the mutual relationships between nodes, construct a social network graph to represent it, and input different social network graphs into the graph neural network model to obtain the node embedding representation of the graph; Step 4, calculating the distance matrix of the node embedding representation of two different social network graphs, and predicting the node pairs of the user alignment task.

[0021] Specifically, refer to Figure 1 The cross-social network user alignment method based on large language model for data enhancement proposed in the present invention comprises the following steps:

[0022] (1) Expansion of user account raw data based on large language model: The original data of user accounts is analyzed by applying the GLM-4-9B large language model with designed prompt words, and the output of the large language model is used as the expanded user feature data;

[0023] (2) Encoding using a fine-tuned pre-trained language model: Fine-tune the pre-trained language model DeBERTa V3, input the original data and expanded user feature data in step (1) into the fine-tuned pre-trained language model DeBERTa V3 for encoding, and obtain the text-based embedded representation vector (i.e., embedded vector) of each user;

[0024] (3) Use the graph neural network model to obtain the embedded representation of the nodes of the graph: Train the graph neural network model. The graph neural network model uses the graph attention model GAT. The embedded representation vector of each user obtained in step (2) is used as the feature vector (node ​​information) of the node, and the relationship between user accounts (follow, friends, etc.) is used as the relationship between nodes to construct a user node graph (as a social network graph). The user node graph is input into the trained graph neural network model to obtain the node embedding vector of each node.

[0025] (4) Matching graph node pairs: Use the node embedding vectors obtained after processing the original data of user accounts on different social networks in steps (1), (2), and (3) to construct a distance matrix, and infer the matching aligned user node pairs based on the distance matrix.

[0026] In step (1), the designed prompt words are used to enhance the original data of the user account using a large language model. The main purpose is to expand the limited user information and find the unchanging deep information behind the variable surface data of different social network accounts of the same natural person. The original data of the user account includes the following: <1> User account attributes: user name, gender, birthday, industry, education background, and personal profile. Due to the design of different platforms and the user's personal filling preferences, account attributes are often missing. <2> Content posted by users: text data of tweets posted by users in the past; the output of the large language model refers to the request made to the large language model through prompt words, and the answer given by the large language model after analyzing the original data of the user account, which includes a data type obtained through discrimination and gives the basis for the discrimination, all expressed in the form of text.

[0027] The content posted by the same user on different platforms is often very different, which makes it difficult to directly compare the original text posted by the user for user alignment. However, these variable contents depend to a large extent on the stable cognitive characteristics of a natural person. This step applies the large language model combined with cognitive psychology theory, and selects relatively mature personality theory, ability theory and language theory to analyze the original data of user accounts. Among them, the personality theory selects the Myers-Briggs Type Indicator (MBTI) for analysis. The Myers-Briggs Type Indicator is a personality type theory model that divides personality into four groups of opposite innate preferences (introversion and extroversion, sensing and intuition, thinking and emotion, judgment and perception, and the four preferences can form 16 stable personality types). The ability theory selects the explicit representation of occupation for analysis. The occupational classification is based on the American Standard Occupational Classification System (SOC), which divides 1016 social occupations into four levels: Major Group, Minor Group, Broad Occupation, and SOCitle. It can clearly distinguish the work content, experience requirements, skill types, professional fields and other characteristics of different occupations. Linguistic theory characterizes and analyzes deep cognitive features by analyzing the dialects and variant grammars of social network users.

[0028] refer to Figure 2 In cognitive psychology personality theory, a person's personality will affect the expression style of his tweets on social platforms. For example, personality types with strong rational tendencies (such as INTP, ENTJ, etc.) may be more inclined to express their opinions directly and objectively, and use less emotional words. Personality types with strong emotional tendencies (such as ENFP, INFP, etc.) may be better at using emotional language to express their feelings and opinions, making the content of the speech more contagious. Extroverted personalities (such as ESTP, ENFP) may be better at incorporating humorous elements into their speeches to make the communication atmosphere more relaxed and pleasant. Introverted or judging personalities (such as INTJ, ISTJ) may pay more attention to the accuracy and depth of the content, and their speaking style is relatively serious; personality will also affect their emotional expression on social platforms. For example, extroverted or emotional personalities (such as ENFP, ESFJ) may be more willing to share their emotional experiences and inner world on social platforms and seek resonance with others. Introverted or thinking personalities (such as INTP, INFJ) may be relatively restrained and rarely express personal emotions in public.

[0029] In cognitive psychology language theory, English dialects and variants around the world (such as British English, American English, Australian English, New Zealand English, Singaporean English, Indian English, etc.) often have their own unique vocabulary and grammar. For example, in Indian English, a mixture of Hindi and English (Hinglish) often appears, such as "I have hazaarthingson my mind right now", where "hazaar" is a Hindi word meaning "one thousand", and the entire sentence adopts the English sentence structure; Australian English contains many unique local words, which reflect the geographical, climate, flora and fauna and cultural characteristics of Australia. "Thongs" refers to flip-flops in Australia, "barbie" refers to a barbecue grill, and "macca's" is a nickname for McDonald's. Some English words have changed in meaning or usage in Australian English. For example, "barbecue" in Australia not only refers to barbecue activities, but is also often used to refer to barbecue grills; "g'day" is a way of greeting, meaning "hello". By identifying these unique language features, the language variants a person habitually uses can be directly reflected in the content they post on social platforms.

[0030] In cognitive psychology ability theory, occupation is a prominent representation of ability. A person's occupation has an impact on the professionalism and relevance of the content they publish. People of different occupations have different professional knowledge and skills in their fields, which will be reflected in their speeches on social platforms. For example, doctors may share health knowledge on the platform, while lawyers may discuss legal cases and opinions. This professionalism makes their speeches more authoritative and credible. Occupational background also determines which topics individuals are more likely to pay attention to and discuss on social platforms. For example, scientific and technological workers may pay more attention to the development of new technologies and new products, while educators may pay more attention to educational policies, teaching methods, etc.; occupational background also affects the position and perspective of the content they publish. For example, environmental workers may be more inclined to express their views on a policy or event from an environmental protection perspective, while entrepreneurs may pay more attention to the impact of the policy or event on the business environment.

[0031] Therefore, this step uses a large language model to conduct an in-depth analysis of the original data of social network user accounts, expands the original data and user attributes into three dimensions: personality characteristics, ability characteristics, and language characteristics, and generates account hybrid characteristics. The specific approach is to design prompt words to query the large language model. The specific design is as follows:

[0032] <1> Analysis of personality traits:

[0033] In terms of personality traits, the Myers-Briggs Type Indicator (MBTI) is used for analysis, and the large language model is used to analyze the personality type based on the text content posted by the user. The design prompt words are as follows:

[0034] Here are some posts from a user posting on the internet.Please inferhis / her MBTI type based on these posts.Your response should clearly provideone MBTI label and give reasons for each dimension in MBTI,posts:'...'

[0035] This prompt specifies the analysis of a specific MBTI personality type and clearly gives the requirements for the specific analysis of the four preference dimensions.

[0036] <2> Analysis of language features:

[0037] In terms of language, a large language model is used to analyze the English variants and dialects used in the text content posted by users to mine the language features behind their text information (including special words, phrases, and grammar). The design prompt words are as follows:

[0038] Here are some posts from a user posting on the internet.Please inferwhat variety of the English language he / she use.Your response should clearly provide one variety label and give reasons,if you can't make a clearinference,please give the answer "not sure",posts:'...'

[0039] The prompt specifies the English variant to be analyzed for the language of a user's posted content, gives the analysis requirements, and requires that an "uncertain" answer be given if there are no clear features. This is because the text posted by the user does not always carry clear features of the variant language.

[0040] <3> Analysis of ability characteristics-occupation:

[0041] In terms of ability characteristics, a large language model is used to analyze the occupational category of users based on the text content they post; the occupational category is based on the US Standard Occupational Classification System (SOC). The design prompt words are as follows:

[0042] Here are some posts from a user posting on the internet.Please inferhis / her careers.Your response should clearly provide one career title inStandard Occupational Classification(SOC)and give reasons,ifyou can't make a clear inference,please give the answer "not sure",posts:'...'

[0043] The prompt word indicates the requirement to analyze the possible occupational classification of a user from the text content posted by a social network and give the reasons, and requires to give an "uncertain" answer if there are no clear characteristics. This is because the occupation of the user cannot always be clearly reflected in the text posted by the user.

[0044] The fine-tuning of the pre-trained language model DeBERTaV3 in step (2) and the training of the graph attention model GAT in step (3) are implemented using the same training method, that is, using the contrastive learning method to make the samples of the same user closer in the encoding space, while the samples of different users are farther away. The specific training method is as follows:

[0045] <1> Forward propagation: Input data into the model to get the model output. The pre-trained language model DeBERTaV3 obtains the word embedding vector of each user account, and the graph attention model GAT obtains the (graph) node embedding vector of each user account.

[0046] <2> Calculate the loss function: The info-NCE loss function is used as the loss function for fine-tuning the model using the contrastive learning method. The specific formula is as follows:

[0047]

[0048]

[0049] Among them, l iis the loss value of a single sample, L is the average loss value of a small batch of samples, M is the total number of small batch samples, z i and are the i-th user account sample in a social network and the corresponding account sample belonging to the same user in another social network, P pos is the set of user accounts, τ is an adjustable hyperparameter;

[0050] <3> Back propagation: Calculate the gradient of the loss value to the model parameters according to the info-NCE loss function, and use the Adam optimization algorithm to update the model parameters according to the loss value;

[0051] <4> Repeat iterations: Repeat the above forward propagation, loss function calculation and back propagation process until the model reaches the specified number of iterations.

[0052] In step (4), the pairwise distance is calculated for the node embedding vectors (node ​​feature vectors) of all users in the two social networks after being processed by steps (1), (2), and (3). The pairwise distance between node embedding vectors is calculated using the cosine distance. The specific formula is as follows:

[0053]

[0054] Among them, u1 and u2 are the feature vectors of two user nodes respectively.

[0055] After calculating the pairwise distances of the node embedding vectors of all user nodes in the two social networks, we get a distance matrix D of size N×M, where N and M are the number of user nodes in the two social networks, respectively. For each user node to be predicted, the pairwise distances of the node pairs formed by it and all nodes in the other social network are sorted, and the node with the smallest pairwise distance (the node in the other social network) is used as the predicted alignment node. The user node to be predicted and the corresponding alignment node form an aligned user node pair as the alignment account result.

[0056] It can be seen that the technical solution of the present invention is divided into the steps of cognitive feature expansion, encoding using a pre-trained language model, obtaining node embedding representation using a graph neural network, matching graph node pairs, etc., involving the following key technologies:

[0057] 1) Cognitive feature expansion

[0058] Due to the differences in functions, community atmosphere and habits of different social platforms, the content displayed by the same natural person user on different social platforms is often very different, but the cognitive characteristics of the user behind the account, such as personality characteristics, language characteristics, ability characteristics, etc., are often fixed. Therefore, the context understanding ability and semantic reasoning ability of the large language model can be used to analyze the deep cognitive characteristics of social account users and obtain extended features of the original account data.

[0059] 2) Encoding using a pre-trained language model

[0060] For the original data and extended data of social network accounts represented in text form, the RoBERTaV3 pre-trained language model is used to convert the text into a fixed-size embedding vector to encode the semantic information in the text into the vector space for easy use in subsequent tasks.

[0061] 3) Use graph neural network to obtain node embedding representation

[0062] In the task of aligning users across social networks, not only the features of the user nodes themselves, but also the features of the users associated with the user nodes can provide features for analysis. Therefore, the graph neural network model is used to encode the node features associated with each node into the embedded representation of the node by utilizing its message passing mechanism between nodes.

[0063] 4) Matching graph node pairs

[0064] The distance between the node embedding vectors of the two graphs is calculated node by node. After obtaining the distance matrix, the feature vector distance of each node to be predicted and the matching node in the other network are sorted to predict the aligned user node.

[0065] The following examples specifically illustrate the cross-social network user alignment method based on hybrid extended cognitive features of the present invention.

[0066] The data used in this embodiment are: 2,000 user account data collected from social platform A and social platform B respectively, which can be connected based on mutual attention relationships, of which 100 user data can be aligned with the other platform. The data examples are as follows:

[0067]

[0068]

[0069] 1. Use a large language model for data enhancement (extension): Ask the large model questions using the three prompt words mentioned above, and use the answers as supplements to obtain an extended data set. The extended data examples are shown in the following table:

[0070]

[0071] 2. Model training: Divide the data into training set and validation set in a ratio of 8:2, train the DeBERTaV3 language model and GAT graph neural network model by comparative learning, and select the model parameters with the best effect on the validation set in 5000 rounds of training as the final model parameters.

[0072] 3. Model reasoning: Input 500 user data of each of the two social networks to be identified into the large language model GLM-4-9B to obtain the expanded data (data format is text); use the expanded data as the input of the fine-tuned DeBERTaV3 language model, and output it as a word embedding vector (data format is a floating point array); use the word embedding vector and the user relationship adjacency matrix (data format is a two-dimensional Boolean array) as the input of the graph neural network, and output it as the graph node embedding vector of each user (data format is a floating point array).

[0073] Identification of aligned user accounts: Calculate the cosine distance between the graph node embedding vectors of each user obtained from the two networks to obtain a distance matrix of size 500×500 (the data format is a two-dimensional floating point matrix). For the target user account to be identified, select all the distance values ​​on the corresponding rows or columns on the distance matrix and sort them. The user node corresponding to the shortest distance value is the aligned user node of the target user account on the other network.

[0074] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for user alignment across social networks based on hybrid extended cognitive features, characterized in that: The following steps are involved: (1) Expansion of user account raw data based on large language model: Design prompt words, and use the designed prompt words to apply the GLM-4-9B large language model to analyze the raw data of user accounts, and use the output of the large language model as the expanded user feature data; (2) Encoding using a fine-tuned pre-trained language model: Fine-tune the pre-trained language model DeBERTaV3, input the original data and expanded user feature data in step (1) into the fine-tuned pre-trained language model DeBERTaV3 for encoding, and obtain the text-based embedding representation vector of each user; (3) Use the graph neural network model to obtain the embedded representation of the nodes of the graph: Train the graph neural network model. The graph neural network model uses the graph attention model GAT. The embedded representation vector of each user obtained in step (2) is used as the feature vector of the node, and the relationship between user accounts is used as the relationship between nodes. The user node graph is constructed and the user node graph is input into the trained graph neural network model to obtain the node embedding vector of each node. (4) Matching graph node pairs: Use the node embedding vectors obtained after processing the original data of user accounts on different social networks in steps (1), (2), and (3) to construct a distance matrix, and infer the matching aligned user node pairs based on the distance matrix.

2. The method according to claim 1, characterized in that In step (1), the original data of the user account is analyzed by applying the GLM-4-9B large language model using the designed prompt words, which can enhance the original data of the user account, expand the limited user information, and find the unchanging deep information behind the variable surface data of different social network accounts of the same natural person. The original data of the user account includes the following: <1> User account attributes: user name, gender, birthday, industry, education background, and personal profile; <2> Content posted by users: text data of tweets posted by users in the past; the output of the large language model refers to the request made to the large language model through prompt words, and the answer given by the large language model after analyzing the original data of the user account, which includes a data type obtained through discrimination and gives the basis for the discrimination, all expressed in the form of text.

3. The method according to claim 1, characterized in that In step (1), the large language model is combined with cognitive psychology theory, and personality theory, ability theory and language theory are selected to analyze the original data of user accounts; among them, the personality theory selects the Myers-Briggs Type Indicator MBTI for analysis; the ability theory selects the explicit representation of occupation for analysis; the language theory analyzes the dialects and variant grammars of social network users to characterize and analyze the deep cognitive characteristics.

4. The method according to claim 3, characterized in that In step (1), when applying the large language model in combination with cognitive psychology theory, and selecting personality theory, ability theory and language theory to analyze the original data of the user account, the original data and user attributes are expanded in three dimensions: personality characteristics, ability characteristics and language characteristics, and the account hybrid characteristics are generated; specifically, prompt words are designed to query the large language model, and the specific design is as follows: <1> Analysis of personality traits: In terms of personality traits, the Myers-Briggs Type Indicator (MBTI) is used for analysis, and the large language model is used to analyze the personality type of users based on the text content they post. The prompts in the design indicate that a specific MBTI personality type needs to be analyzed, and require reasons for the specific analysis of the four preference dimensions; <2> Analysis of language features: In terms of language, a large language model is used to analyze the English variants and dialects used in the text content posted by users to explore the language features behind their text information; The prompts are designed to ask for analysis of the English variant to which a user's post belongs, and to give reasons for the analysis, and to give an "unsure" answer if there is no clear feature; <3> Analysis of capability characteristics: In terms of ability characteristics, a large language model is used to analyze the user's occupational category based on the text content posted by the user; the designed prompt words require analysis of the possible occupational classification of a user from the text content posted on a user's social network and give reasons, and require an "uncertain" answer if there are no clear characteristics.

5. The method according to claim 1, characterized in that The fine-tuning of the pre-trained language model DeBERTa V3 in step (2) and the training of the graph attention model GAT in step (3) are implemented using the same training method, that is, using the contrastive learning method, so that the samples of the same user are closer in the encoding space, while the samples of different users are farther away.

6. The method according to claim 5, characterized in that The following training method is used to fine-tune the pre-trained language model DeBERTaV3 in step (2) and train the graph attention model GAT in step (3): <1> Forward propagation: Input data into the model to get the model output. The pre-trained language model DeBERTaV3 obtains the word embedding vector of each user account, and the graph attention model GAT obtains the node embedding vector of each user account. <2> Calculate the loss function: The info-NCE loss function is used as the loss function for fine-tuning the model using the contrastive learning method. The specific formula is as follows: Among them, l i is the loss value of a single sample, L is the average loss value of a small batch of samples, M is the total number of small batch samples, z i and are the i-th user account sample in a social network and the corresponding account sample belonging to the same user in another social network, P pos is the set of user accounts corresponding to the user pair, and τ is an adjustable hyperparameter; <3> Back propagation: Calculate the gradient of the loss value to the model parameters according to the info-NCE loss function, and use the Adam optimization algorithm to update the model parameters according to the loss value; <4> Repeat iterations: Repeat the above forward propagation, loss function calculation and back propagation process until the model reaches the specified number of iterations.

7. The method according to claim 1, characterized in that In step (4), the pairwise distances are calculated for the node embedding vectors of all users in the two social networks after being processed by steps (1), (2), and (3). The pairwise distances between node embedding vectors are calculated using the cosine distance. The specific formula is as follows: Among them, u1 and u2 are the feature vectors of two user nodes respectively; After calculating the pairwise distances of the node embedding vectors of all user nodes in the two social networks, a distance matrix D of size N×M is obtained, where N and M are the number of user nodes in the two social networks respectively; for each user node to be predicted, the pairwise distances of its node pairs composed of all nodes in the other social network are sorted, and the node with the smallest pairwise distance is used as the predicted alignment node. The user node to be predicted and the corresponding alignment node form an aligned user node pair.

8. The method according to any one of claims 1 to 7, characterized in that This method is applied in the fields of artificial intelligence and social network data mining technology.

9. A system for implementing the method according to any one of claims 1 to 7.

10. The system according to claim 9, characterized in that The system is used in the fields of artificial intelligence and social network data mining technology.

Citation Information

Patent Citations

  • Cross-social network user identity recognition method and system based on machine learning

    CN109753602A

  • Cross-network user alignment method based on deep learning

    CN110347932A

  • Multi-source heterogeneous network user alignment method based on graph neural network

    CN113095948A

  • Social network alignment method and system based on multi-information fusion and graph optimization

    CN116776008A