User identity alignment method, device, equipment and medium integrating stance analysis
By introducing a position analysis module into the identity alignment model, identifying the user's position and calculating the similarity between users, the problem of user identity alignment between different social platforms is solved, and the recognition accuracy and the effect of intelligent recommendation services are improved.
Patent Information
- Application Number
- CN202211579541.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-09
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-12-09
AI Technical Summary
How to effectively align user identities between different social platforms to determine whether users in different platforms belong to the same natural person.
By obtaining user feature data in multiple social networks, inputting them into the pre-trained identity alignment model, using the position analysis module to identify the user's position, and computing the similarity between users. If the similarity is greater than the preset threshold, it is determined that multiple users are the same natural person.
It improves the accuracy of the identification of the user identity alignment model, can more accurately align user identity, and improves the effectiveness of applications such as intelligent recommendation services.
Smart Images

Figure CN116167885B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of social network analysis, and in particular to a user identity alignment method, device, equipment and medium integrating stance analysis. Background Art
[0002] With the widespread development and widespread use of social media, more and more natural persons have begun to register virtual identities on different social platforms to enjoy the different services provided by the platforms. For example, Weibo provides users with timely news, which users can forward and comment on, and users can also publish their own articles and diaries, as well as form fans and interest groups; Douban provides a wealth of books and audio-visual ratings for enthusiasts to choose and learn from, and allows users to form interest groups to discuss in them, and users can also organize various offline activities.
[0003] Due to the diversity of social platforms and the flexibility of user identities, how to determine that users on different platforms belong to the same natural person, that is, to align user identities across platforms, is a technical problem that needs to be urgently solved by technical personnel in this field. Summary of the invention
[0004] The embodiments of the present application provide a method, apparatus, device and medium for aligning user identities in a fusion stance analysis. In order to have a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary is not a general review, nor is it intended to identify key / important components or describe the scope of protection of these embodiments. Its only purpose is to present some concepts in a simple form as a preface to the detailed description that follows.
[0005] In a first aspect, an embodiment of the present application provides a user identity alignment method integrating stance analysis, including:
[0006] Obtain characteristic data of users in multiple social networks;
[0007] Inputting the feature data into a pre-trained identity alignment model to obtain similarities between users in different social networks; wherein the identity alignment model includes a stance analysis module for identifying the stance of the user based on the feature data;
[0008] If the similarity between the users is greater than a preset threshold, it is determined that the multiple users are the same natural person.
[0009] In an optional embodiment, obtaining characteristic data of users in multiple social networks includes:
[0010] Obtain name data, region data, gender data, affiliation data, social network structure data, and user posting data of users in multiple social networks.
[0011] In an optional embodiment, before inputting the user's feature data into the pre-trained identity alignment model, the method further includes:
[0012] constructing the identity alignment model;
[0013] The identity alignment model includes a multi-layer attribute embedding module, a vector fusion module, and a similarity calculation module that are sequentially connected.
[0014] In an optional embodiment, the multi-layer attribute embedding module includes:
[0015] A name feature vector extraction unit, used to pre-process the user's name data through a preset word vector model to obtain a name feature vector;
[0016] An attribute feature vector extraction unit, used to pre-process the user's region data, gender data, and affiliated institution data through a preset word embedding model to obtain an attribute feature vector;
[0017] A network structure feature vector extraction unit is used to pre-process the user's social network structure data through a preset graph representation learning algorithm to obtain a network structure feature vector; wherein the network structure includes a user attention relationship graph, a like relationship graph, or a forwarding relationship graph;
[0018] The topic stance feature vector extraction unit is used to perform topic analysis on the user's published content data through a preset topic model to obtain a topic feature vector; to perform stance analysis on the user's published content data through a pre-trained stance analysis model to obtain a stance feature vector corresponding to the topic; and to multiply the topic feature vector and the corresponding stance feature vector to obtain the user's topic stance feature vector.
[0019] In an optional embodiment, before inputting the user's feature data into the pre-trained identity alignment model, the method further includes:
[0020] Constructing a dataset for text stance analysis;
[0021] Adding stance labels to the data set, wherein the stance labels include positive stance, negative stance, and neutral stance;
[0022] The classification model is trained based on the labeled data set to obtain a trained stance analysis model.
[0023] In an optional embodiment, the vector fusion module is used to fuse the name feature vector, attribute feature vector, network structure feature vector and topic stance feature vector in each social network to obtain the user feature vector in each social network.
[0024] In an optional embodiment, the similarity calculation module is used to map the user feature vectors in each social network to the same vector space through a preset vector mapping algorithm; and calculate the similarity between the user feature vectors in different social networks in the same vector space.
[0025] In a second aspect, an embodiment of the present application provides a user identity alignment device integrating stance analysis, including:
[0026] An acquisition module, used to acquire characteristic data of users in multiple social networks;
[0027] An analysis module, used to input the feature data into a pre-trained identity alignment model to obtain the similarity between users in different social networks; wherein the identity alignment model includes a stance analysis module, used to identify the user's stance based on the feature data;
[0028] The identification module is used to determine that multiple users are the same natural person if the similarity between the users is greater than a preset threshold.
[0029] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory storing program instructions, wherein the processor is configured to execute the user identity alignment method for fusion stance analysis provided in the above embodiment when executing the program instructions.
[0030] In a fourth aspect, an embodiment of the present application provides a computer-readable medium having computer-readable instructions stored thereon, and the computer-readable instructions are executed by a processor to implement a user identity alignment method for integrating stance analysis provided in the above embodiment.
[0031] The technical solution provided by the embodiments of the present application may have the following beneficial effects:
[0032] The user identity alignment method integrated with stance analysis provided in the embodiment of the present application inputs the acquired user features into the identity alignment model. The identity alignment model introduces a stance analysis module, which can perform stance analysis on user behavior. Stance, as an abstract feature, cannot be set or changed artificially, and can make the user portrait more three-dimensional and rich, closer to the natural person characteristics in the real world. The recognition accuracy of the user identity alignment model can be effectively improved. At the same time, users with the same focus topic but different stances can also be distinguished. By accurately aligning user identities, the effects of subsequent applications such as intelligent recommendation services can be improved.
[0033] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0035] Figure 1 is a flowchart of a method for aligning user identities by integrating stance analysis according to an exemplary embodiment;
[0036] Figure 2 is a schematic diagram showing the representation and combination of a feature vector according to an exemplary embodiment;
[0037] Figure 3 is a structural diagram of an identity alignment model according to an exemplary embodiment;
[0038] Figure 4 is a structural schematic diagram of a user identity alignment device integrating stance analysis according to an exemplary embodiment;
[0039] Figure 5 is a schematic structural diagram of an electronic device according to an exemplary embodiment;
[0040] Figure 6 It is a schematic diagram of a computer storage medium according to an exemplary embodiment. DETAILED DESCRIPTION
[0041] The following description and the drawings sufficiently illustrate specific embodiments of the invention to enable those skilled in the art to practice them.
[0042] It should be clear that the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0043] When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are only examples of systems and methods consistent with some aspects of the present invention as detailed in the attached claims.
[0044] Social networks play an increasingly important role in modern society, and each natural person has formed different virtual user identities on different social platforms. User alignment refers to finding out whether two users on different social networks are the same natural person in the real world. This task is of great significance for platforms to provide personalized services, solve the cold start problem of websites, and combat cybercrime.
[0045] At present, the research on user identity alignment generally has a shallow level of exploration of user behavior. Most of them only analyze the surface features of user-generated content, such as time information and trajectory information, but do not consider the deep features of user-generated content, such as user opinions and positions.
[0046] On social platforms such as Weibo and Douban, users often express opinions with certain stances in areas of their interest. So, similar to gender, topic, etc., stance should also be used as a feature to characterize user portraits to identify natural persons.
[0047] First, stance reflects a person's attitude towards something. Like the topics that users follow, stance is often long-lasting and difficult to change. Even in different social networks, users' stances are generally the same. Second, unlike user settings such as usernames and regions, stance is an abstract feature that cannot be set or changed manually. Finally, stance is often based on the topic, which can make the user portrait more three-dimensional and rich, and closer to the characteristics of natural people in the real world.
[0048] Based on this, the embodiment of the present application provides a user identity alignment method integrating stance analysis. Figure 1 The user identity alignment method for integrating stance analysis provided in the embodiment of the present application is introduced in detail. Figure 1 , the method specifically includes the following steps.
[0049] S101 obtains characteristic data of users in multiple social networks.
[0050] In a possible implementation, user identities of multiple social networking platforms may be analyzed, for example, Weibo, Douban, Zhihu, Xiaohongshu, Twitter, Meta, etc.
[0051] Furthermore, characteristic data of users in each social network to be analyzed is obtained. Specifically, the user's name data is obtained, and the user's attribute data can also be obtained from the user's homepage, and the attribute data includes region data, gender data, affiliated institution data, etc., such as which school the user belongs to, which unit the user belongs to, user profile, user occupation, etc.
[0052] Furthermore, the user's social network structure data is obtained, such as the user's forwarding relationship data, like relationship data, follow relationship data, and user published content data. The embodiment of the present application not only focuses on the user's follow relationship, but also analyzes the user's like and forwarding relationship data.
[0053] S102 inputs the feature data into a pre-trained identity alignment model to obtain the similarity between users in different social networks; wherein the identity alignment model includes a stance analysis module for identifying the user's stance based on the feature data.
[0054] In one embodiment, before inputting the user's feature data into the pre-trained identity alignment model, the method further includes: constructing an identity alignment model, wherein the identity alignment model includes a sequentially connected multi-layer attribute embedding module, a vector fusion module, and a similarity calculation module.
[0055] Specifically, the user's feature data is first input into the multi-layer attribute embedding module, which can be mainly divided into:
[0056] Name features: These are composed of the characters in the user name and are the most basic user attributes.
[0057] Attribute characteristics: The basic unit is the words or phrases in the user's basic attribute data (region, organization);
[0058] Network structure characteristics: mainly refers to the structural characteristics of users' attention relationship network, forwarding relationship network, and like relationship network;
[0059] Topic and stance features: that is, user behavior, mainly the user topics and stances obtained after subject and stance analysis on articles or posts published or forwarded by users.
[0060] For example, user 'Tommy Li' belongs to Peking University and published an article titled "The research of useridentity linkage by latent user space". Then 'Tommy Li' is the user's name feature, 'Beijing University' is the user's attribute feature, and the vector obtained after thematic analysis of the paper title "The research of useridentity linkage by latent user space" is the user's topic feature. You can also perform stance analysis on the content to obtain the stance features corresponding to the user and the topic, and form a topic stance feature vector based on the topic features and stance features. It can also include network structure features, mainly referring to the user's attention relationship network, forwarding relationship network, like relationship network and other structural features.
[0061] In one embodiment, the multi-layer attribute embedding module includes: a name feature vector extraction unit, an attribute feature vector extraction unit, a network structure feature vector extraction unit, and a topic stance feature vector extraction unit.
[0062] Among them, the name feature vector extraction unit mainly extracts vectors from the user's name data through a preset word vector model to obtain a name feature vector.
[0063] The attribute feature vector extraction unit mainly pre-processes the user's attribute data, such as region data, gender data, and affiliated institution data, through a preset word embedding model to obtain an attribute feature vector.
[0064] In an optional embodiment, the user's regional data, gender data, affiliated institution data and other attribute data are preprocessed using the Word2Vec algorithm to obtain an attribute feature vector.
[0065] In one embodiment, the word-level features in the user attributes are trained using the Word2Vec model. The Word2Vec algorithm converts text into feature vectors. i The word-level attributes of Indicates that it contains phrases or short sentences such as gender, region, affiliation, educational experience, etc. can be divided into a sequence of m words w i =w i1 ,w i2 ,…,w ik ,…w im , where w ik yes The expression of the kth word in .
[0066] To give a more detailed example: in the user "Yoshua Bengio", its word-level attributes are Department of Computer Science and Operations Research, QC, Canada", which will be vectorized into a sequence of words followed by the word count: w i = "University": 1, "of": 2, "Montr'eal": 1, "Department": 1, ..., others: 0. Then, each word is represented by a unique number as an ID.
[0067] use to represent the word-level attribute set of a social network with n users. Each w i Both are used for word-level documents in deep learning.
[0068] Optionally, other word embedding models such as TF-IDF and GloVe can also be used.
[0069] The network structure feature vector extraction unit is used to pre-process the user's social network structure data through a preset graph representation learning algorithm to obtain a network structure vector.
[0070] In an optional embodiment, a user's attention relationship graph, like relationship graph, or forwarding relationship graph is obtained, which can be a directed graph or an undirected graph, and the network structure is extracted by the LINE graph representation algorithm. The LINE graph representation algorithm is a scalable network node representation method that represents the entire network by obtaining local structure and global structure. The LINE algorithm uses BFS (breadth-first algorithm) to construct a neighborhood and can be applied to weighted graphs. The specific approach is: the local structure (first-order similarity) is obtained through the similarity between adjacent nodes, and the global structure (second-order similarity) is obtained through the similarity between nodes with the same neighbors. Finally, the two similarity-trained vectors are spliced into the entire network structure.
[0071] Optionally, you can also use Deepwalk Node2Vec, GCN (Graph Convolutional Network), GAT (Graph attention networks) and other graph representation algorithms.
[0072] The topic stance feature vector extraction unit is used to perform topic analysis on the user's published content data through a preset topic model to obtain a topic feature vector; to perform stance analysis on the user's published content data through a pre-trained stance analysis model to obtain a stance feature vector corresponding to the topic; and to multiply the topic feature vector and the corresponding stance feature vector to obtain the user's topic stance feature vector.
[0073] Specifically, the method provided in the embodiment of the present application not only analyzes the subject information of the content posted by the user, but also analyzes the user's position based on the subject data.
[0074] In an optional embodiment, the LDA (Latent Dirichlet Allocation) algorithm is an unsupervised topic analysis model, which is essentially a three-layer Bayesian model. LDA considers documents to be the probability distribution of topics, and topics to be the probability distribution of vocabulary. The number of topics is artificially given, and then documents are input in batches for training. The output obtained includes two: one is the vocabulary probability distribution of each topic, and the other is the probability distribution of the topic in each article. The embodiment of the present application first uses the LDA topic analysis model to perform topic analysis on user published content to obtain a topic feature vector. Optionally, other topic models can also be used.
[0075] Furthermore, a stance analysis model is trained to perform stance analysis on the content posted by the user to obtain a second behavior vector.
[0076] In one embodiment, a data set for text stance analysis is constructed; stance labels are added to the data set, wherein the stance labels include positive stance, negative stance, and neutral stance; and a classification model is trained based on the labeled training data set to obtain a trained stance analysis model.
[0077] In one possible implementation, a network dataset related to user-published content, such as restaurant reviews, is obtained as a training dataset. The data in the dataset mainly consists of online users' reviews and star ratings of a restaurant, which is also a form of review that is often seen on review software. The reviews contain many emotional words such as "very average", "like", and "delicious", which are very suitable corpus for sentiment or stance analysis. Therefore, using it to train a stance analysis model can also help improve the operating efficiency of the model.
[0078] Before training the model, the data needs to be preprocessed, including word segmentation and converting the star rating to classification results for stance analysis. The purpose of word segmentation is to convert feature text into word vectors for input into the model; converting the star rating to classification is because we need to limit the number of labels in the model to two, which is also in line with the goal of stance analysis. Therefore, a function is artificially designed to convert the star rating in the data into stance, and it is considered that the star rating greater than 4 is a positive evaluation, represented by 1; the star rating less than or equal to 2 is a negative evaluation, represented by -1, and 3 and 4 are neutral evaluations, represented by 0. The annotation results are added as data labels.
[0079] The classification model is trained according to the training data set after adding labels to obtain a trained stance analysis model, which can give a stance evaluation for any text. Among them, the classification model adopted in the embodiment of the present application can be a naive Bayes model. The goal of stance analysis is to divide the user's stance into three types: positive, negative, and neutral. The naive Bayes model belongs to a classification model, which matches the goal of stance analysis. The naive Bayes model is a classification model based on prior probability, supported by classical mathematical theory, and the model effect has a relatively high stability. Compared with other classification models, the naive Bayes model has a higher tolerance for data integrity. When the data set is incomplete or missing, the naive Bayes model can still play an ideal classification role.
[0080] Optionally, those skilled in the art may also adopt other classification models, such as support vector machines, logistic regression and other classification models, which are not specifically limited in the embodiments of the present application.
[0081] After obtaining the trained stance analysis model, the trained stance analysis model is used to perform stance analysis on the user's published content to obtain a stance feature vector corresponding to the topic; the topic feature vector is multiplied by the corresponding stance feature vector to obtain the user's topic stance feature vector. The topic stance feature vector contains both the topic information published by the user and the user's stance information on the topic, which can describe the character's characteristics at a deeper level.
[0082] In a possible implementation, the identity alignment model also includes a vector fusion module for fusing the user name feature vector, attribute feature vector, network structure feature vector, and topic stance feature vector in each social network to obtain the user feature vector in each social network.
[0083] In a possible implementation, when fusing multiple vectors, a feature vector concatenation method may be used, for example, a weighted concatenation method may be used to fuse multiple feature vectors, or a weighted summation method may be used to fuse multiple feature vectors. The specific weights may be obtained through attention mechanism training.
[0084] like Figure 2 As shown in Figure 1, stance analysis and topic analysis are performed in one module, collectively referred to as the user behavior module. First, topic analysis is performed on the content published by the user, and stance analysis is performed based on the topic analysis. The matrices obtained by the two analysis algorithms are multiplied to obtain the final user behavior feature vector.
[0085] exist Figure 2 In the vector fusion shown, each component of the topic analysis matrix is multiplied by the one-dimensional component of the corresponding stance analysis matrix to obtain a new user behavior vector emb_t.
[0086] For the user name, the user's region and the user's friend relationship, the corresponding algorithms are used to convert them into vectors respectively: for the user name, a simple word vectorization model is used to obtain the character-level vector emb_c; for the user's region as a word-level feature, the Word2Vec algorithm is used for vector representation to obtain the word-level vector emb_w; for the user's friend relationship as the network topology, the graph embedding algorithm LINE is used for vector representation to obtain the network structure vector emb_s.
[0087] After obtaining the individual vector representation of each feature, the four feature vectors are fused to obtain the fused multi-level attribute vector embed, that is, Figure 2 The fusion vector Y in .
[0088] In a possible implementation, the identity alignment model also includes a similarity calculation module, which is used to map the user feature vectors in each social network to the same vector space through a preset vector mapping algorithm; and calculate the similarity between the user feature vectors in different social networks in the same vector space.
[0089] Specifically, the RCCA (Regularized Canonical Correlation Analysis) vector mapping algorithm can be used to convert the feature vectors of different networks into the same vector space using seed user pairs. The optimization goal of the conversion process is to make the vector representations of the seed user pairs as close as possible. Among them, RCCA works well in maximizing the distance between linear projection vectors. Therefore, the embodiment of the present application adopts the RCCA method to study the identity connection between similar users in the user alignment problem.
[0090] Most RCCA algorithms define the typical matrix as H = [h1,h2,...h i ,...h k ]∈R d*k and M=[m1,m2,...m j ,...m k ]∈R d*k , including k pairs of linear projections. The typical matrix of user alignment probabilities is obtained by a series of carefully designed linear projections to transform the relevant social network G x / G y The corresponding vectors X / Y are mapped to a common correlation space Z to solve H T X and M T The correlation ρ between Y is maximized. Then, G X and G Y The identity connection between potential identical users can be estimated by comparing the distances of their vectorized features in Z.
[0091] Furthermore, the similarity between vectors is calculated. The Euclidean distance between the feature vectors of users in different social networks can be calculated, and the similarity between users in different social networks can be obtained based on the Euclidean distance. The cosine similarity between the feature vectors of users in different social networks can also be calculated, and the similarity between users in different social networks can be obtained based on the cosine similarity between the feature vectors.
[0092] In an exemplary scenario, the RCCA algorithm is used to map the vectors corresponding to the two social networks into the same vector space Z. The corresponding vectors of X and Y in the same vector space Z are Z x and Z yThe similarity between nodes is calculated by comparing the Euclidean distance D in the vector space Z. The smaller the spatial distance D between the vectors corresponding to two nodes in the vector space Z, the higher the possibility that the two nodes are the same natural person.
[0093] Figure 3 is a schematic diagram of a structure of an identity alignment model according to an exemplary embodiment. Figure 3 In the embodiment shown, for two social networks G x and G y ,First, their user behaviors (content published by users) and other user attributes (user name, organization, region, network topology such as forwarding relationships between users, etc.).
[0094] The user behavior is analyzed for topics and positions to obtain the embedded behavior-level vector; the user relationship is embedded as a network structure to obtain a structure-level vector; other user attributes (user name, organization, region) are also embedded in the corresponding vector to obtain a character-level name vector and a word-level attribute vector. The four vectors are fused to obtain two fused vectors X and Y (social network G x Corresponding vector X, social network G y corresponding vector Y).
[0095] Furthermore, using the RCCA algorithm, the vectors corresponding to the two social networks are mapped to the same vector space Z. The corresponding vectors of X and Y in vector space Z are Z x and Z y The similarity between nodes is calculated by comparing the Euclidean distance D in the vector space Z. The smaller the spatial distance D between the vectors corresponding to two nodes in the vector space Z, the higher the possibility that the two nodes are the same natural person. The similarity between different users can be obtained.
[0096] According to this step, by building an identity alignment model, the user's feature data is analyzed to obtain the similarity between users. When calculating the similarity between users, the user's stance features are introduced on the basis of common features such as user name, region, and topic. By adding stance features, the character portrait can be portrayed in more depth and the accuracy of the analysis can be improved.
[0097] S103: If the similarity between the users is greater than a preset threshold, it is determined that the multiple users are the same natural person.
[0098] In a possible implementation, after obtaining the similarity between users, it is determined whether the similarity between users is greater than a preset threshold. If it is greater than the preset threshold, it is determined that multiple users in different social networks are the same natural person. If it is less than or equal to the preset threshold, it is determined that multiple users in different social networks may not be the same natural person.
[0099] Identifying the same natural person on different platforms is conducive to network service providers providing more targeted personalized services. For example, the goods and services followed by the same user on different social platforms are related. If the social platform can obtain the goods or topics that a newly registered user likes on other social platforms, it can provide personalized recommendation services when there is no local data of the user; in addition, user identity alignment can also combat false and untrue information and illegal activities such as online fraud on the Internet from the source. Before the criminal gangs on multiple social platforms are matched, cybercrime can only be cracked down on after the fact, and it is impossible to prevent it in advance. If the accounts of the same criminal gang on different platforms can be matched, crimes can be effectively prevented before they commit crimes.
[0100] The embodiment of the present application proposes a user identity alignment model that integrates stance analysis. The model adds stance analysis on the basis of topic analysis of user behavior, and integrates network structure features and user basic attribute features for vector embedding. At the same time, regular correlation analysis (RCCA) is used to project user vectors into the same vector space, and the user similarity is calculated by comparing the vector distance in the same vector space, so as to match user pairs.
[0101] Experimental results show that stance analysis of user behavior can effectively improve the accuracy of the user alignment model. At the same time, it can also distinguish users who focus on the same topic but have different stances.
[0102] The present application also provides a user identity alignment device for fusion stance analysis, which is used to execute the user identity alignment method for fusion stance analysis of the above embodiment, such as Figure 4 As shown, the device comprises:
[0103] An acquisition module 401 is used to acquire characteristic data of users in multiple social networks;
[0104] An analysis module 402 is used to input the feature data into a pre-trained identity alignment model to obtain similarities between users in different social networks; wherein the identity alignment model includes a stance analysis module for identifying the stance of the user based on the feature data;
[0105] The identification module 403 is used to determine that multiple users are the same natural person if the similarity between the users is greater than a preset threshold.
[0106] It should be noted that the user identity alignment device for fusion stance analysis provided in the above embodiment only uses the division of the above functional modules as an example when executing the user identity alignment method for fusion stance analysis. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the user identity alignment device for fusion stance analysis provided in the above embodiment and the user identity alignment method for fusion stance analysis belong to the same concept, and the implementation process thereof is detailed in the method embodiment, which will not be repeated here.
[0107] The embodiment of the present application also provides an electronic device corresponding to the user identity alignment method for integrating stance analysis provided in the above embodiment, so as to execute the user identity alignment method for integrating stance analysis.
[0108] Please refer to Figure 5 , which shows a schematic diagram of an electronic device provided by some embodiments of the present application. Figure 5 As shown, the electronic device includes: a processor 500, a memory 501, a bus 502 and a communication interface 503, and the processor 500, the communication interface 503 and the memory 501 are connected via the bus 502; the memory 501 stores a computer program that can be run on the processor 500, and when the processor 500 runs the computer program, the user identity alignment method for fusion stance analysis provided in any of the aforementioned embodiments of the present application is executed.
[0109] The memory 501 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 503 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used.
[0110] The bus 502 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The memory 501 is used to store programs, and the processor 500 executes the program after receiving the execution instruction. The user identity alignment method for fusion stance analysis disclosed in any of the embodiments of the present application may be applied to the processor 500, or implemented by the processor 500.
[0111] The processor 500 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 500. The above processor 500 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor to be executed, or the hardware and software modules in the decoding processor can be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 501, and the processor 500 reads the information in the memory 501 and completes the steps of the above method in combination with its hardware.
[0112] The electronic device provided in the embodiment of the present application and the user identity alignment method for fusion stance analysis provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented therein.
[0113] The present application also provides a computer-readable storage medium corresponding to the user identity alignment method for integrating stance analysis provided in the above embodiment. Figure 6 The computer-readable storage medium shown is a CD 600 on which a computer program (ie, a program product) is stored. When the computer program is executed by a processor, the user identity alignment method for fusion stance analysis provided in any of the aforementioned embodiments is executed.
[0114] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not listed here one by one.
[0115] The computer-readable storage medium provided in the above-mentioned embodiment of the present application and the user identity alignment method for fusion stance analysis provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by the application program stored therein.
[0116] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0117] The above embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the attached claims.
Claims
1. A user identity alignment method integrating stance analysis, characterized in that: include: Acquire characteristic data of users in multiple social networks; wherein the characteristic data includes name characteristics, attribute characteristics, network structure characteristics, and topic stance characteristics; before inputting the characteristic data of the user into the pre-trained identity alignment model, it also includes: constructing the identity alignment model; the identity alignment model includes a sequentially connected multi-layer attribute embedding module, a vector fusion module, and a similarity calculation module; wherein the multi-layer attribute embedding module includes: a name feature vector extraction unit, which is used to pre-process the user's name data through a preset word vector model to obtain a name feature vector; an attribute feature vector extraction unit, which is used to pre-process the user's region data, gender data, and affiliated data through a preset word embedding model. The system is used to pre-process the social network structure data of the user to obtain the attribute feature vector; the network structure feature vector extraction unit is used to pre-process the social network structure data of the user through a preset graph representation learning algorithm to obtain the network structure feature vector; wherein the network structure includes a user attention relationship graph, a like relationship graph or a forwarding relationship graph; the subject position feature vector extraction unit is used to perform subject analysis on the user's published content data through a preset subject model to obtain a subject feature vector; and is used to perform position analysis on the user's published content data through a pre-trained position analysis model to obtain a position feature vector corresponding to the subject; and the subject feature vector is multiplied by the corresponding position feature vector to obtain the user's subject position feature vector; Inputting the feature data into a pre-trained identity alignment model to obtain similarities between users in different social networks; wherein the identity alignment model includes a stance analysis module for identifying the stance of the user based on the feature data; If the similarity between the users is greater than a preset threshold, it is determined that the multiple users are the same natural person.
2. The method according to claim 1, characterized in that Get user feature data from multiple social networks, including: Obtain name data, region data, gender data, affiliation data, social network structure data, and user posting data of users in multiple social networks.
3. The method according to any one of claims 1 to 2, characterized in that: Before inputting the user's feature data into the pre-trained identity alignment model, the method further includes: Constructing a dataset for text stance analysis; Adding stance labels to the data set, wherein the stance labels include positive stance, negative stance, and neutral stance; The classification model is trained based on the labeled data set to obtain a trained stance analysis model.
4. The method according to claim 1, characterized in that The vector fusion module is used to fuse the name feature vector, attribute feature vector, network structure feature vector and topic stance feature vector in each social network to obtain the user feature vector in each social network.
5. The method according to claim 4, characterized in that The similarity calculation module is used to map the user feature vectors in each social network to the same vector space through a preset vector mapping algorithm; and calculate the similarity between the user feature vectors in different social networks in the same vector space.
6. A user identity alignment device integrating stance analysis, characterized in that: include: An acquisition module, used to acquire characteristic data of users in multiple social networks; wherein the characteristic data includes name characteristics, attribute characteristics, network structure characteristics, and subject position characteristics; The analysis module is used to input the feature data into a pre-trained identity alignment model to obtain the similarity between users in different social networks; wherein the identity alignment model includes a stance analysis module, which is used to identify the user's stance according to the feature data; before the feature data of the user is input into the pre-trained identity alignment model, it also includes: constructing the identity alignment model; the identity alignment model includes a multi-layer attribute embedding module, a vector fusion module, and a similarity calculation module connected in sequence; wherein the multi-layer attribute embedding module includes: a name feature vector extraction unit, which is used to pre-process the user's name data through a preset word vector model to obtain a name feature vector; an attribute feature vector extraction unit, which is used to pre-process the user's name data through a preset word embedding model to obtain a name feature vector; The user's regional data, gender data, and affiliated institution data are preprocessed to obtain an attribute feature vector; a network structure feature vector extraction unit is used to preprocess the user's social network structure data through a preset graph representation learning algorithm to obtain a network structure feature vector; wherein the network structure includes a user attention relationship graph, a like relationship graph, or a forwarding relationship graph; a topic stance feature vector extraction unit is used to perform topic analysis on the user's published content data through a preset topic model to obtain a topic feature vector; and is used to perform stance analysis on the user's published content data through a pre-trained stance analysis model to obtain a stance feature vector corresponding to the topic; and the topic feature vector is multiplied by the corresponding stance feature vector to obtain the user's topic stance feature vector; The identification module is used to determine that multiple users are the same natural person if the similarity between the users is greater than a preset threshold.
7. An electronic device, characterized in that: The invention comprises a processor and a memory storing program instructions, wherein the processor is configured to execute the user identity alignment method for fusion stance analysis according to any one of claims 1 to 5 when executing the program instructions.
8. A computer-readable medium, characterized in that Computer-readable instructions are stored thereon, and the computer-readable instructions are executed by a processor to implement a user identity alignment method integrating stance analysis as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Cross-space target virtual identity association method based on multilayer attribute analysis
CN113779520A
Multi-feature fusion cross-social network user identity association method
CN114581254A
Cross-social network user identity matching method based on entropy weight method, medium and device
CN115048563A