A method and device for identifying a social media key debunker
By constructing a three-layer echo chamber network structure, key debunkers on social media are identified, solving the problem of neglecting the propagation network structure and user interaction behavior in existing technologies. This enables effective dissemination of debunking information and ensures rapid suppression of rumors.
Patent Information
- Application Number
- CN202310168452.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-17
- Filing Date
- 2023-02-24
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2043-02-24
AI Technical Summary
Existing technologies neglect the relationship between the propagation network structure and user interaction behavior in social media, making it difficult to effectively identify key debunkers and affecting the dissemination of debunking information.
By constructing a three-layer echo chamber network structure, including a user network, an event network, and an echo chamber network, and combining sentiment features, influence features, and relationship features, key debunker IDs are screened out, and target user IDs with large sentiment features that match the preset sentiment feature, influence feature value, and relationship feature value are identified and screened out.
It enables the effective identification of key debunkers on social media, allowing for rapid and timely guidance of public opinion, suppression of rumor spread, and improved debunking effectiveness.
Smart Images

Figure CN116127202B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of computer science and technology, and in particular to a method and device for identifying key debunkers of social media. BACKGROUND
[0002] Debunking refers to publishing clarifying information about the content of rumors to reduce the negative impact of rumors. The key to improving the effect of debunking is to make people no longer believe rumors. According to existing literature, the factors affecting the effect of debunking can be summarized as follows: topic, timing, channel and content. In previous studies, people have considered the effect of rumor refutation. However, these studies ignore the individual differences in the response of information recipients to rumor refutation and the different community backgrounds of the disseminators. Some studies have shown that individuals with different knowledge bases, interests and values have different perceived credibility when refuting rumors. In addition, social identity plays a key role in the dissemination of rumor refutation. Individuals form social identity by establishing emotional connections with social organizations or other social groups, and feel a sense of belonging to groups of individuals who share their beliefs. This social identity is accompanied by a confirmation bias, which makes them tend to share views, information or opinions with the same values, and rarely accept other views, which ultimately leads to the formation of an echo chamber.
[0003] However, existing research on rumor refutation often ignores the relationship between the structure of the dissemination network and the behavior of user interaction. In the process of dissemination of rumor refutation information, some users of social media play a key role in influencing the behavior of other users and guiding the direction of public opinion, and these users are called key debunkers. Therefore, identifying these key debunkers is crucial for the dissemination of debunking information.
[0004] How to identify key debunkers in social media is a problem to be solved at present. SUMMARY
[0005] The present application provides a method and device for identifying key debunkers, which are used for identifying key debunkers in an echo chamber.
[0006] The first aspect of the present application provides a method for identifying key debunkers, comprising:
[0007] obtaining original debunking texts, and obtaining user IDs that forward and / or comment on the original debunking texts;
[0008] identifying and classifying text topics of the original debunking texts;
[0009] identifying a set of echo chambers corresponding to each class of text topics, wherein an echo chamber in the set of echo chambers includes at least two target user IDs, and the target user IDs are user IDs of users participating in commenting or forwarding at least two same original debunking texts in the text topics.
[0010] constructing a user network according to the user IDs, the user network being a network with the user IDs as nodes and interaction relationships between different user IDs as edges;
[0011] constructing an event network according to the original debunking texts, the event network including a set of original debunking texts, a set of derivative texts of the original debunking texts, a set of text topics of the original debunking texts, and a set of topics corresponding to the set of text topics;
[0012] constructing an echo chamber network according to the echo chambers, the echo chamber network being a network with the echo chambers as nodes and common original debunking texts between different echo chambers as edges;
[0013] constructing a three-layer echo chamber network structure according to the user network, the event network, and the echo chamber network, the echo chamber network structure being a network structure in which the user network, the event network, and the echo chamber network are associated with each other;
[0014] constructing a key debunker identification model according to the echo chamber network structure, and screening key debunker IDs in the echo chambers through the key debunker identification model, the key debunker ID being a target user ID with a sentiment feature meeting a preset sentiment feature, an influence feature value greater than a preset influence feature value, and a relationship feature value greater than a preset relationship feature value.
[0015] Optionally, the echo chamber set corresponding to each type of text topic is identified, the echo chambers in the echo chamber set including at least two target user IDs, the target user ID being a user ID participating in commenting or forwarding at least two same original debunking texts in the text topic, and the target user ID including:
[0016] obtaining the classified at least two types of text topics;
[0017] filling at least two original debunking texts under the first type of text topic into a first queue;
[0018] extracting a first original debunking text from the first queue, determining a first user ID set corresponding to the first original debunking text, the first user ID set being a set of user IDs forwarding and / or commenting on the first original debunking text;
[0019] respectively matching the first user ID set with other user ID sets corresponding to at least one other original debunking text in the first queue until the first user ID set is matched with all the other user ID sets corresponding to the other original debunking texts in the first queue, and obtaining a plurality of first intersections of the first user ID set and the other user ID sets;
[0020] respectively match a second user ID set of a second original rumor-busting text with other user ID sets corresponding to at least one other original rumor-busting text in the first queue until the second user ID set is matched with all other user ID sets corresponding to all other original rumor-busting texts in the first queue, obtain a plurality of second intersections of the second user ID set and the other user ID sets;
[0021] determine a plurality of first echo chamber sets corresponding to the first text topic according to the first intersection and the second intersection;
[0022] fill original rumor-busting texts under a second text topic into a second queue, and determine a plurality of second echo chamber sets corresponding to the second text topic.
[0023] Optionally, the filtering of the key rumor-buster ID in the echo chamber by the key rumor-buster identification model comprises:
[0024] filtering, by the key rumor-buster identification model, a first target user ID in the target user ID in the echo chamber, the emotional feature of the first target user ID meeting a preset emotional feature;
[0025] filtering, by the key rumor-buster identification model, a second target user ID in the first target user ID, the influence feature of the second target user ID being greater than a preset influence feature;
[0026] filtering, by the key rumor-buster identification model, a key rumor-buster ID in the second target user ID, the relationship feature of the key rumor-buster ID meeting a preset relationship feature.
[0027] Optionally, the filtering of the first target user ID in the target user ID in the echo chamber by the key rumor-buster identification model, the emotional feature of the first target user ID meeting a preset emotional feature comprises:
[0028] obtaining a derivative text of the original rumor-busting text commented by the target user ID;
[0029] analyzing, by an emotional tendency analysis interface, an emotional polarity of the derivative text corresponding to the target user ID, the emotional polarity comprising a positive polarity and a negative polarity;
[0030] determining whether there is at least one derivative text with a negative polarity in the derivative text of the target user ID;
[0031] if yes, isolating the target user ID;
[0032] if no, determining the target user ID as the first target user ID.
[0033] Optionally, the second target user ID in the first target user ID is screened through the key debunking person identification model, and an influence feature value of the second target user ID is greater than a preset influence feature value;
[0034] The number of echo chambers in which the first target user ID participates is identified through the echo chamber network;
[0035] And a user activity coefficient of the first target user ID is determined through a first formula according to the number of echo chambers, the first formula being:
[0036]
[0037] Wherein, H u is the user activity coefficient of the first target user ID, P is the number of all text topics, ECS is the number of all echo chambers, EC u is the number of echo chambers in which the first target user ID participates, p u is the number of text topics in which the first target user ID participates;
[0038] A user creativity coefficient of the first target user ID is determined according to a second formula;
[0039] The second formula is:
[0040] Wherein, C u is the user creativity coefficient of the first target user ID, N u is the number of original debunking texts that the first target user ID forwards or publishes in a target time period, and h is the target time period;
[0041] A user text quality coefficient of the first target user ID is determined according to the number of original debunking texts that the first target user ID forwards or publishes in a target time period and a third formula;
[0042] The third formula is:
[0043] Wherein, Q u is the user text quality coefficient, N u is the number of original debunking texts that the first target user ID forwards or publishes in a target time period, L u is the forwarding quantity of original debunking texts that the first target user ID forwards or publishes in a target time period, Q u is the comment quantity of original debunking texts that the first target user ID forwards or publishes in a target time period, and Y uThe number of likes of the original rumor-busting text or the published rumor-busting text is forwarded for the first target user ID in a target time period;
[0044] The user content quality coefficient of the first target user ID is determined according to a fourth formula through the user creativity coefficient and the user text quality coefficient;
[0045] The fourth formula is:
[0046] M u = C u × Q u ;
[0047] Wherein, M u is the user content quality coefficient of the first target user ID, C u is the user creativity coefficient of the first target user ID, and Q u is the user text quality coefficient of the first target user ID;
[0048] The influence characteristic value of the first target user ID is determined according to the user activity coefficient and the user content quality coefficient of the first target user ID through a fifth formula;
[0049] The fifth formula is:
[0050]
[0051] Wherein, A u is the influence characteristic value of the first target user ID, H u is the user activity coefficient of the first target user ID, M u is the user content quality coefficient of the first target user ID, MAXH u is the maximum user activity coefficient in the first target user ID, and MAXM u is the maximum user content quality coefficient in the first target user ID;
[0052] It is judged whether the influence characteristic value of the first target user ID is greater than the preset influence characteristic value;
[0053] If yes, the first target user ID is determined as the second target user ID.
[0054] Optionally, the key rumor-buster ID in the second target user ID is screened through the key rumor-buster identification model, and the relationship characteristic value of the key rumor-buster ID is greater than a preset relationship characteristic value, which includes:
[0055] The user event prestige coefficient of the second target user ID is determined according to a sixth formula;
[0056] The sixth formula is:
[0057]
[0058] wherein, I u is the user event prestige coefficient of the second target user ID, R u is the number of all original debunking texts published and / or forwarded by the second target user ID in the target time period being forwarded and / or commented on by other users, Z u is the number of all original debunking texts published and / or forwarded by the second target user ID being liked by other users, h is the target time period;
[0059] determine the user participation coefficient of the second target user ID according to a seventh formula;
[0060] The seventh formula is:
[0061]
[0062] wherein, F u is the user participation coefficient of the second target user ID, t u is the number of original debunking texts commented on and / or forwarded by the second target user ID, T is the number of all original debunking texts;
[0063] determine the relationship feature value of the second target user ID according to an eighth formula through the user event prestige coefficient and the user participation coefficient;
[0064] The eighth formula is:
[0065]
[0066] wherein, D u is the relationship feature value of the second target user ID, I u is the user event prestige coefficient of the second target user ID, F u is the user participation coefficient of the second target user ID, maxI u is the maximum value of the user event prestige coefficient of the second target user ID, maxF u is the maximum value of the user participation coefficient of the second target user ID;
[0067] determine whether the relationship feature value of the second target user ID is greater than a preset relationship feature value;
[0068] If yes, determine that the second target user ID is a key debunking person ID.
[0069] Optionally, after the echo chamber network structure is constructed according to the key debunking person identification model, and the key debunking person ID of the echo chamber is screened through the key debunking person identification model, the identification method further comprises:
[0070] The effectiveness of the key debunking person identification model is verified based on the key debunking person ID.
[0071] Optionally, before the text theme of the original debunking text is identified and the text theme is classified, the identification method further comprises:
[0072] The original debunking text is preprocessed, and the preprocessing comprises filtering repeated and / or invalid original debunking texts;
[0073] The identification and classification of the text theme of the original debunking text comprises:
[0074] The text theme of the original debunking text after preprocessing is identified and classified.
[0075] Optionally, the identification and classification of the text theme of the original debunking text comprises:
[0076] The text theme of the original debunking text is identified and classified through a LAD theme space model, and the LAD theme space model has a "document-theme-word" three-layer generative Bayesian network structure.
[0077] The second aspect of the present application provides an identification device of a social media key debunking person, comprising:
[0078] An acquisition unit is configured to acquire original debunking texts and acquire user IDs forwarding and / or commenting on the original debunking texts;
[0079] A first identification unit is configured to identify and classify text themes of the original debunking texts;
[0080] A second identification unit is configured to identify a set of echo chambers corresponding to each type of text theme, wherein the echo chambers in the set of echo chambers comprise at least two target user IDs, and the target user IDs are user IDs participating in commenting or forwarding at least two same original debunking texts in the text theme;
[0081] A first construction unit is configured to construct a user network according to the user IDs, wherein the user network is a network taking the user IDs as nodes and taking interactive relationships between different user IDs as edges;
[0082] a second constructing unit configured to construct an event network according to the original rumor-busting texts, the event network comprising a set of original rumor-busting texts, a set of derivative texts of the original rumor-busting texts, a set of text topics of the original rumor-busting texts, and a set of topics corresponding to the set of text topics;
[0083] a third constructing unit configured to construct an echo chamber network according to the echo chambers, the echo chamber network being a network with the echo chambers as nodes and common original rumor-busting texts between different echo chambers as edges;
[0084] a fourth constructing unit configured to construct a three-layer echo chamber network structure according to the user network, the event network, and the echo chamber network, the echo chamber network structure being a network structure in which the user network, the event network, and the echo chamber network are associated with each other;
[0085] a screening unit configured to construct a key rumor-buster identification model according to the echo chamber network structure, and screen a key rumor-buster ID in the echo chambers through the key rumor-buster identification model, the key rumor-buster ID being a target user ID with an emotional feature meeting a preset emotional feature, an influence feature value greater than a preset influence feature value, and a relationship feature value greater than a preset relationship feature value.
[0086] Optionally, the second identifying unit is specifically configured to:
[0087] obtain the at least two categories of text topics after classification;
[0088] fill at least two original rumor-busting texts under a first category of text topics into a first queue;
[0089] extract a first original rumor-busting text from the first queue, determine a first user ID set corresponding to the first original rumor-busting text, the first user ID set being a set of user IDs forwarding and / or commenting on the first original rumor-busting text;
[0090] respectively match the first user ID set with other user ID sets corresponding to at least one other original rumor-busting text in the first queue until the first user ID set is matched with all the other user ID sets corresponding to all the other original rumor-busting texts in the first queue, and obtain a plurality of first intersections of the first user ID set and the other user ID sets;
[0091] respectively match a second user ID set of a second original rumor-busting text with other user ID sets corresponding to at least one other original rumor-busting text in the first queue until the second user ID set is matched with all the other user ID sets corresponding to all the other original rumor-busting texts in the first queue, and obtain a plurality of second intersections of the second user ID set and the other user ID sets.
[0092] determining a plurality of first echo chamber sets corresponding to the first type of text subject according to the first intersection and the second intersection;
[0093] filling original debunking text under a second type of text subject into a second queue, and determining a plurality of second echo chamber sets corresponding to the second type of text subject.
[0094] Optionally, the screening unit is specifically configured to:
[0095] screening a first target user ID in the target user ID in the echo chamber through the key debunking person identification model, and the emotional feature of the first target user ID meets a preset emotional feature;
[0096] screening a second target user ID in the first target user ID through the key debunking person identification model, and the influence feature of the second target user ID is greater than a preset influence feature;
[0097] screening a key debunking person ID in the second target user ID through the key debunking person identification model, and the relationship feature of the key debunking person ID meets a preset relationship feature.
[0098] Optionally, the screening unit is specifically configured to:
[0099] obtaining derivative text of the original debunking text commented by the target user ID;
[0100] analyzing the emotional polarity of the derivative text corresponding to the target user ID through an emotional tendency analysis interface, the emotional polarity including positive polarity and negative polarity;
[0101] determining whether there is at least one piece of derivative text with negative polarity in the derivative text of the target user ID;
[0102] if yes, isolating the target user ID;
[0103] if no, determining the target user ID as the first target user ID.
[0104] Optionally, the screening unit is specifically configured to:
[0105] identifying the number of echo chambers participated by the first target user ID through the echo chamber network;
[0106] and determining the user activity coefficient of the first target user ID according to the number of echo chambers through a first formula, the first formula being:
[0107]
[0108] wherein, Hu is a user activity coefficient of the first target user ID, P is a number of all text topics, ECS is a number of all echo chambers, EC u is a number of echo chambers participated by the first target user ID, p u is a number of text topics participated by the first target user ID;
[0109] determining a user creativity coefficient of the first target user ID according to a second formula;
[0110] The second formula is:
[0111] wherein, C u is a user creativity coefficient of the first target user ID, N u is a number of original rumor-busting texts retweeted or published by the first target user ID in a target time period, h is the target time period;
[0112] determining a user text quality coefficient of the first target user ID according to a number of original rumor-busting texts retweeted or published by the first target user ID in a target time period and a third formula;
[0113] The third formula is:
[0114] wherein, Q u is a user text quality coefficient, N u is a number of original rumor-busting texts retweeted or published by the first target user ID in a target time period, L u is a number of retweets of original rumor-busting texts retweeted or published by the first target user ID in a target time period, Q u is a number of comments of original rumor-busting texts retweeted or published by the first target user ID in a target time period, Y u is a number of likes of original rumor-busting texts retweeted or published by the first target user ID in a target time period;
[0115] determining a user content quality coefficient according to a fourth formula through the user creativity coefficient and the user text quality coefficient;
[0116] The fourth formula is:
[0117] M u =C u ×Q u ;
[0118] wherein, M u is a user content quality coefficient of the first target user ID, C ucreate a user creativity coefficient Q for a user with a first target user ID u create a user content quality coefficient for the user with the first target user ID
[0119] determine an influence characteristic value of the first target user ID according to the user activity coefficient and the user content quality coefficient of the first target user ID by a fifth formula
[0120] The fifth formula is:
[0121]
[0122] wherein A u is the influence characteristic value of the first target user ID, H u is the user activity coefficient of the first target user ID, M u is the user content quality coefficient of the first target user ID, MAXH u is the maximum user activity coefficient in the first target user ID, MAXM u is the maximum user content quality coefficient in the first target user ID
[0123] determine whether the influence characteristic value of the first target user ID is greater than the preset influence characteristic value
[0124] if yes, determine the first target user ID as a second target user ID
[0125] Optionally, the screening unit is specifically configured to:
[0126] determine a user event prestige coefficient of the second target user ID according to a sixth formula
[0127] The sixth formula is:
[0128]
[0129] wherein I u is the user event prestige coefficient of the second target user ID, R u is the number of all original rumor-busting texts published and / or forwarded by the second target user ID in a target time period and forwarded and / or commented by other users, Z u is the number of all original rumor-busting texts published and / or forwarded by the second target user ID and liked by other users, h is the target time period
[0130] determine a user participation coefficient of the second target user ID according to a seventh formula
[0131] The seventh formula is:
[0132]
[0133] wherein, F u is a user participation coefficient of the second target user ID, t u is a number of original debunking texts in which the second target user ID participates in forwarding and / or commenting, and T is a number of all original debunking texts;
[0134] The relationship feature value of the second target user ID is determined according to an eighth formula by using the user event prestige coefficient and the user participation coefficient.
[0135] The eighth formula is as follows:
[0136]
[0137] wherein, D u is a relationship feature value of the second target user ID, I u is a user event prestige coefficient of the second target user ID, F u is a user participation coefficient of the second target user ID, maxI u is a maximum value of the user event prestige coefficient of the second target user ID, maxF u is a maximum value of the user participation coefficient of the second target user ID.
[0138] It is determined whether the relationship feature value of the second target user ID is greater than a preset relationship feature value.
[0139] If yes, the second target user ID is determined as a key debunking person ID.
[0140] Optionally, the identification device further comprises:
[0141] A verification unit is configured to verify the effectiveness of the key debunking person identification model based on the key debunking person ID.
[0142] Optionally, the identification device further comprises:
[0143] A preprocessing unit is configured to preprocess the original debunking texts, and the preprocessing includes filtering duplicate and / or invalid original debunking texts.
[0144] The first identification unit is specifically configured to:
[0145] Identify and classify the text topics of the original debunking texts after preprocessing.
[0146] Optionally, the first identification unit is specifically configured to:
[0147] identify a text theme of the original rumor-busting text through a LAD theme space model, and classify the text theme, the LAD theme space model having a "document-theme-word" three-layer generative Bayesian network structure.
[0148] The third aspect of the present application provides a device for identifying a key rumor-buster in social media, the device comprising:
[0149] a processor, a memory, an input / output unit, and a bus;
[0150] the processor being connected to the memory, the input / output unit, and the bus;
[0151] the memory storing a program, and the processor invoking the program to execute the method for identifying a key rumor-buster in social media according to the first aspect and any one of the optional aspects of the first aspect.
[0152] The fourth aspect of the present application provides a computer-readable storage medium storing a program, the program being executed on a computer to execute the method for identifying a key rumor-buster in social media according to the first aspect and any one of the optional aspects of the first aspect.
[0153] As can be seen from the above technical solutions, the present application has the following advantages: obtaining original rumor-busting texts and user IDs that forward and / or comment on the original rumor-busting texts; identifying text themes of the original rumor-busting texts, classifying the text themes, identifying echo chamber sets corresponding to each class of text themes, the echo chamber set including at least two echo chambers, the echo chamber including at least two target user IDs, the target user ID being a user ID that participates in commenting or forwarding texts in at least two same original rumor-busting texts; constructing a user network according to the user IDs; constructing an event network according to the original rumor-busting texts; constructing an echo chamber network according to the echo chambers; constructing a three-layer echo chamber network structure according to the user network, the event network, and the echo chamber network; constructing a key rumor-buster identification model according to the echo chamber network structure, and screening key rumor-buster IDs in the echo chambers through the key rumor-buster identification model, the key rumor-buster ID being a target user ID with a sentiment feature meeting a preset sentiment feature, an influence feature value greater than a preset influence feature value, and a relationship feature value greater than a preset relationship feature value; through the above method, a key rumor-buster identification model is constructed, and key rumor-buster IDs in each echo chamber are identified through the key rumor-buster identification model, so that the public opinion direction of each echo chamber can be effectively guided through the key rumor-buster IDs, and rumors can be quickly and timely suppressed. BRIEF DESCRIPTION OF DRAWINGS
[0154] In order to more clearly illustrate the technical solutions in the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.
[0155] Figure 1 An embodiment flowchart of a social media key debunker identification method provided by the present application;
[0156] Figure 2 Another embodiment flowchart of a social media key debunker identification method provided by the present application;
[0157] Figure 3 An embodiment structure diagram of a social media key debunker identification device provided by the present application;
[0158] Figure 4 Another embodiment structure diagram of a social media key debunker identification device provided by the present application;
[0159] Figure 5 An embodiment structure diagram of a social media key debunker identification device provided by the present application. DETAILED DESCRIPTION
[0160] The present application provides a social media key debunker identification method and identification device, which are used to identify key debunkers in an echo chamber.
[0161] It should be noted that the social media key debunker identification method provided by the present application can be applied to a terminal, and can also be applied to a server, for example, the terminal can be a smart phone or a computer, a tablet computer, a smart television, a smart watch, a portable computer terminal, and a fixed terminal such as a desktop computer. For convenience of description, the present application takes the terminal as an execution subject for example.
[0162] Please refer to Figure 1 , Figure 1 An embodiment of a social media key debunker identification method provided by the present application includes:
[0163] 101. The terminal acquires the original debunking text and acquires the user ID of the user who forwards and / or comments on the original debunking text;
[0164] In the embodiment, the terminal uses a Python crawler to capture original rumor-busting texts within a target time range according to rumor-busting related keywords such as "debunking", "clarification", and "rumor refutation" by searching function on a data collection platform, and capture user IDs that forward, comment on, and / or like the original rumor-busting texts; the data collection platform can be Weibo or WeChat public account, etc.
[0165] 102. The terminal identifies and classifies the text theme of the original rumor-busting text;
[0166] In the embodiment, the terminal identifies the text theme of the original rumor-busting text and classifies the text theme by a LAD document theme generation model (Latent Dirichlet Allocation, LAD). The LAD is a topic space model, which is a topic model algorithm based on the idea of probability statistics, has a "document-theme-word" three-layer generative Bayesian network structure, and can be used to identify potential topic information in large-scale events or corpus. The identification of the text theme is only related to the content of the original rumor-busting text and is irrelevant to the comment content of the original rumor-busting text. The text theme can be divided into three categories: social theme, livelihood theme, and entertainment theme. For example, the original rumor-busting text is a message about XX cases on a certain Shoulu and isolation, which attracts the attention and concern of netizens. After verification, the district has not received a report on the discovery of XX cases, and the message is a rumor. The terminal identifies and classifies the text theme of the original rumor-busting text as a social theme, and the text topic corresponding to the text theme of the original rumor-busting text is XX debunking.
[0167] 103. The terminal identifies a sound studio set corresponding to each type of text theme, wherein the sound studio set in the sound studio set includes at least two target user IDs, and the target user ID is a user ID participating in commenting or forwarding at least two same original rumor-busting texts in the text theme;
[0168] In this embodiment, a text topic can correspond to several echo rooms. Each echo room includes at least two target user IDs, which are user IDs who participated in commenting on or forwarding at least two identical original debunking text topics within that text topic. For example: a social text topic contains three original debunking texts: 1. Original debunking text for the XX incident; 2. Original debunking text for the XX assault incident; 3. Original debunking text for the XX soy sauce additive incident; User ID A forwarded the original debunking texts for both the XX incident and the XX assault incident; User ID B forwarded the original debunking texts for both the XX assault incident and the XX soy sauce additive incident; User ID C forwarded the original debunking texts for the XX incident, the XX assault incident, and the XX soy sauce additive incident. In this case, the social topic corresponds to two echo rooms. The first echo room includes User ID A and User ID C; the second echo room includes User ID B and User ID C.
[0169] 104. The terminal constructs a user network based on the user ID. This user network is a network built with user IDs as nodes and the interaction relationships between different user IDs as edges.
[0170] In this embodiment, the user network is a network used to study the evolution of collaborative relationships between different user IDs. The user network is a network built with user IDs as nodes and the interaction relationships between different user IDs as edges; the model for constructing the user network is as follows: M u =(U,E u-u ,W(E u-u )); where, U={u1, u2,…,u n} is a set of user IDs; E u-u ={(u i ,u j )|θ(u i ,u j )}, i, j=1,2,...,n; among them, if θ(u i ,u j If θ(u) = 0, it means there is no interaction relationship between the two user IDs; if θ(u) = 0, it means there is no interaction relationship between the two user IDs. i ,u j If ) = 1, it indicates that there is an interaction relationship between the two user IDs; the weight of the edge in the user network is W(E). u-u )={w(u i ,u j )|u i ,u j ∈U}; where the weights of the edges in the user network are used to represent the degree of interaction between user IDs, for example: user ID u i With user IDu jForwarding or commenting on the same original debunking text counts as one interaction, and the user IDu i With user IDu j Forwarding or commenting on two identical original debunking texts constitutes two interactions.
[0171] 105. The terminal constructs an event network based on the original debunking text. The event network includes the original debunking text set, the derived text set of the original debunking text, the text topic set of the original debunking text, and the topic set corresponding to the text topic set.
[0172] In this embodiment, if the original debunking text set is defined as T, then the original debunking text set T = {t0, t1, t2, ..., t} n}, where t0, t1, t2, ..., t n The original debunking text is the source text; the derived texts are the comment texts following the original debunking text. If the set of derived texts is defined as T... ′ Then the derived text set T ′ ={t ′ 0, t ′ 1,t ′ 2,…,t ′ n If we define the set of text topics as P, then the set of text topics P = {P0, P1, P2, ..., P} n}, where each topic contains at least two original debunking texts; if the set of topics corresponding to the set of text topics is defined as B, then the topic set B = {b0, b1, b2, ..., b} n}, where b i This is a set of topics corresponding to each type of text theme.
[0173] 106. The terminal constructs an echo chamber network based on the echo chambers. This echo chamber network is a network constructed with echo chambers as nodes and common original debunking texts between different echo chambers as edges.
[0174] In this embodiment, the terminal constructs an echo chamber network based on the echo chamber, wherein the model for constructing the echo chamber network is as follows: M G =(G,E e-e ,W(E e-e ), where G={g0,g1,g2,…,g n} is a node in the echo chamber network; E e-e E represents the edge between different echo chambers. e-e ={(g i g j )|θ(g i g j )},i,j=1,2,…,n;where, if θ(gi j ) = 0, it indicates that there is no common original rumor-busting text between two echo chambers; if θ(g i j ) = 1, it indicates that there is common original rumor-busting text between two echo chambers; W(E e-e ) is the weight of the edge in the echo chamber network, which represents the number of edges of the common original rumor-busting text between different echo chambers, and the thickness of the edge represents the number of target user IDs participating in the original rumor-busting text, wherein W(E e-e ) = {w(g i , g j ) | g i , g j ∈ G}.
[0175] 107、Terminal constructs a three-layer echo chamber network structure according to the user network, the event network and the echo chamber network, and the echo chamber network structure is a network structure in which the user network, the event network and the echo chamber network are associated with each other;
[0176] In the embodiment, there are complex corresponding relationships between the user and the event, the event and the echo chamber, and the user and the echo chamber, and the user network, the event network and the echo chamber network are connected together to form a three-layer complex network structure in which the user, the event and the echo chamber are associated with each other; the specific association manner is as follows: I. The relationship set between the user network and the event network is E U-T = {(u i , t j ) | α(u i , t j ) = 1}, wherein α(u i , t j ) = 1 indicates that the user ID u j participates in the discussion of the original rumor-busting text t j , and the weight of the mapping relationship edge is W(E U-T ) = {w(u i , t j ) | u i ∈ U, t j ∈ T}, and the weight of the edge represents the number of user IDs participating in the discussion of the original rumor-busting text t j ; there is a mutual mapping relationship between the user ID and the event (original rumor-busting text), and the mutual mapping relationship has two representation manners: one is that represents the number of user IDs participating in the same original rumor-busting text t j ; and the other is that represents the same user ID U i The number of original debunking texts involved; the mapping relationship between the event network and the user network represents the association under the user ID interaction. The relationship set between the echo chamber network and the event network is: E G-T = {(g i , t j ) | β(g i , t j ) = 1}, where β(g i , t j ) = 1 represents the event (original debunking text) t i in the echo chamber g j , and the weight of the mapping relationship edge is: W(E G-T ) = {w(g i , t j ) | g i ∈ G, t j ∈ T}, where the edge weight represents the number of events in the echo chamber g i ; the echo chamber and the event have a mutual mapping relationship, which has two representations: one is , which represents the number of echo chambers participating in the same event (original debunking text) t j ; the other is , which represents the events (original debunking texts) contained in the same echo chamber g i . The mapping relationship set between the echo chamber network and the user network is: E G-U = {(g i , u j ) | χ(g i , u j ) = 1}, where χ(g i , u j ) = 1 represents that the user ID u i is contained in the echo chamber g j , i.e., the user ID u j participates in the echo chamber g i , and the same user ID can participate in different echo chambers; the relationship weight is defined as: W(E G-U ) = {w(g i , u j ) | g i ∈ G, u i ∈ U}, which is used to represent the participation degree of a user ID in the echo chamber network; the mapping relationship between the echo chamber and the user ID has two representations: one is , which represents the number of echo chambers participated in by the user ID u j , which reflects the activity of the user in the social network from this aspect; the other is , which represents the echo chamber g iThe number of user IDs contained in the echo chamber can determine the size of the echo chamber.
[0177] 108、The terminal constructs a key debunker identification model according to the echo chamber network structure, and screens the key debunker ID in the echo chamber through the key debunker identification model, the key debunker ID being a target user ID whose emotional feature meets a preset emotional feature, whose influence feature value is greater than a preset influence feature value, and whose relationship feature value is greater than a preset relationship feature value.
[0178] In this embodiment, the screening of the key debunker ID in the echo chamber through the key debunker identification model includes: screening a first target user ID of the target user ID in the echo chamber through the key debunker identification model, the emotional feature of the first target user ID meeting the preset emotional feature; screening a second target user ID in the first target user ID through the key debunker identification model, the influence feature of the second target user ID being greater than the preset influence feature; and screening a key debunker ID in the second target user ID through the key debunker identification model, the relationship feature of the key debunker ID meeting the preset relationship feature. The specific screening process will be described in detail in the next embodiment, and will not be described here.
[0179] In this embodiment, the key debunker identification model is constructed through the above method, and the key debunker ID in each echo chamber is identified through the key debunker identification model, so that the public opinion direction of each echo chamber can be effectively guided through the key debunker ID, and the rumor can be quickly and timely suppressed.
[0180] To make the key debunker identification method provided by the present application more obvious and easy to understand, the key debunker identification method provided by the present application will be described in detail as follows:
[0181] Please refer to Figure 2 , Figure 2 Another embodiment of the key debunker identification method provided by the present application includes:
[0182] 201、The terminal acquires the original debunking text, and acquires the user IDs forwarding and / or commenting on the original debunking text.
[0183] The step 201 in this embodiment is similar to the step 101 in the foregoing Figure 1 embodiment, and will not be described here.
[0184] 202、The terminal pre-processes the acquired original debunking text, and the pre-processing includes filtering repeated and / or invalid original debunking texts.
[0185] In the embodiment, the terminal pre-processes the obtained original refutation text, including: deleting repeated and / or blank invalid original refutation text through data filtering, and deleting useless symbols in the original refutation text, thereby reducing data noise of the original refutation text; eliminating original refutation text with a vocabulary less than a preset vocabulary, and eliminating original refutation text with a forwarding quantity less than a preset forwarding quantity, thereby improving the representativeness of the sample; deleting interference original refutation text published by water army and / or robot accounts, identifying water army ID and / or robot account ID through a SocialBotHunter algorithm, including: extracting a feature vector according to social behavior of a user ID, calculating an initial outlier value of the user ID according to the feature vector; associating a binary random variable with the user ID, so that social interaction between user IDs is modeled as a two-by-two Markov random field (MRF); transmitting original refutation text published by the user ID into the MRF, and modifying an abnormal score, thereby identifying the water army ID and / or the robot account ID. Through the above pre-processing, the credibility of the pre-processed original refutation text is improved.
[0186] 203. The terminal identifies and classifies the text topics of the pre-processed original refutation text;
[0187] The step 203 in the embodiment is similar to the step 102 in the foregoing embodiment, and details are not repeated here. Figure 1
[0188] 204. The terminal identifies an echo chamber set corresponding to each type of text topic, and the echo chamber in the echo chamber set includes at least two target user IDs, and the target user ID is a user ID participating in commenting or forwarding at least two same original refutation texts in the text topic;
[0189] In the embodiment, the terminal identifying the echo chamber set corresponding to each type of text theme comprises: obtaining at least two types of classified text themes; filling at least two original rumor-busting texts under the first type of text theme into a first queue; extracting a first original rumor-busting text from the first queue, determining a first user ID set corresponding to the first original rumor-busting text, the first user ID set being a user ID set forwarding and / or commenting on the first original rumor-busting text; matching the first user ID set with other user ID sets corresponding to at least one other original rumor-busting text in the first queue until the matching of the first user ID set with the other user ID sets corresponding to all the other original rumor-busting texts in the first queue is completed, obtaining several first intersections of the first user ID set and the other user ID sets; respectively matching a second user ID set of a second original rumor-busting text with the other user ID sets corresponding to at least one other original rumor-busting text in the first queue until the matching of the second user ID set with the other user ID sets corresponding to the other original rumor-busting texts in the first queue is completed, obtaining several second intersections of the second user ID set and the other user ID sets; determining several first echo chamber sets corresponding to the first type of text theme according to the first intersections and the second intersections; and filling original rumor-busting texts under the second type of text theme into a second queue and determining several second echo chamber sets corresponding to the second type of text theme. For example, there are three original rumor-busting texts t1, t2 and t3 under the first type of text theme, t1, t2 and t3 are filled into the first queue, t1 is extracted, and the first user ID set corresponding to t1 is U1={u1, u2, u3, u4}. If the second user ID set corresponding to t2 is U2={u1, u2, u4, u6}, and the third user ID set corresponding to t3 is U3={u1, u3, u4, u5, u6, u7}, the first intersections obtained by the terminal include U 1,2 ={u1, u2, u4}, U 1,3 ={u1, u3, u4}, and U 1,2,3 ={u1, u4}; the second intersections obtained by the terminal include U 2,1 ={u1, u2, u4}, U 2,3 ={u1, u4, u6}, and U 2,1,3 ={u1, u4}; the first echo chamber sets corresponding to the first type of text theme determined by the terminal include G 1,2 ={u1, u2, u4}, G 1,3 ={u1, u3, u4}, G 1,2,3 ={u1, u4}, G 2,1 ={u1, u2, u4}, G 2,3 ={u1, u4, u6}, and G 2,1,3 ={u1, u4}.
[0190] 205、the terminal constructs a user network according to the user ID, the user network being a network with the user ID as a node and an interaction relationship between different user IDs as an edge;
[0191] 206、the terminal constructs an event network according to the original debunking text, the event network including an original debunking text set, a derivative text set of the original debunking text, and a text theme set of the original debunking text, i.e., a topic set corresponding to the text theme set;
[0192] 207、the terminal constructs an echo chamber network according to the echo chamber, the echo chamber network being a network with the echo chamber as a node and a common original debunking text between different echo chambers as an edge;
[0193] 208、the terminal constructs a three-layer echo chamber network structure according to the user network, the event network, and the echo chamber network, the echo chamber network structure being a network structure in which the user network, the event network, and the echo chamber network are associated with each other;
[0194] Steps 205 to 208 in this embodiment are similar to steps 104 to 107 in the foregoing Figure 1 embodiment, and will not be described here in detail.
[0195] 209、the terminal constructs a key debunker identification model according to the echo chamber network structure, and screens a key debunker ID in the echo chamber through the key debunker identification model, the key debunker ID being a target user ID with an emotional feature meeting a preset emotional feature, an influence feature greater than a preset influence feature, and a relationship feature greater than a preset relationship feature;
[0196] In the embodiment, the terminal screens the first target user ID in the target user ID in the echo chamber through the key debunking person identification model, and the emotional feature of the first target user ID meets the preset emotional feature, including: obtaining the diffraction text of the target user ID commenting on the original debunking text, analyzing the emotional polarity of the diffraction text corresponding to the target user ID through an emotional tendency analysis interface, and the emotional polarity includes positive polarity and negative polarity; determining whether there is at least one diffraction text with negative emotional polarity in the diffraction text of the target user ID; if yes, isolating the target user ID; if not, determining the target user ID as the first target user ID; in the embodiment, the emotional polarity of the diffraction text corresponding to the target user ID can be determined by using the emotional tendency analysis interface in the Baidu AI platform, and the diffraction text commented by the target user ID is scored in terms of emotional tendency, and the score of the emotional tendency score is [-1, 1], wherein -1 represents absolute negative emotional polarity, and 1 represents absolute positive emotional polarity; the emotional tendency analysis interface in the Baidu AI platform is for Chinese text with subjective description in general or specific scene, automatically judges the emotional polarity category of the Chinese text, and gives the corresponding confidence, and the emotional polarity is divided into positive polarity and negative polarity. On the basis of the user ID emotional polarity analysis, the application screens the positive first target user ID according to the "two-level transmission theory".
[0197] II. The terminal screens the second target user ID in the first target user ID through the key debunking person identification model, and the influence value of the second target user ID is greater than the preset influence value, including: identifying the number of echo chambers participated by the first target user ID through the echo chamber network; and determining the user activity coefficient of the first target user ID according to the number of echo chambers through a first formula, the first formula is: Wherein, H u is the user activity coefficient of the first target user ID, P is the number of all text topics, ECS is the number of all echo chambers, EC u is the number of echo chambers participated by the first target user ID, p u is the number of text topics participated by the first target user ID; the user creativity coefficient of the first target user ID is determined according to a second formula; the second formula is: Wherein, C u is the user creativity coefficient of the first target user ID, N u is the number of the first target user ID forwarding the original debunking text or publishing the original debunking text in the target time period, and h is the target time period; the user text quality coefficient of the first target user ID is determined according to the number of the first target user ID forwarding the original debunking text or publishing the original debunking text in the first target time period and a third formula, the third formula is: Wherein, Q u is the user text quality coefficient, Nu the number of forwarding original rumor-busting texts or publishing rumor-busting texts of the first target user ID in the target time period, L u the forwarding quantity of forwarding original rumor-busting texts or publishing rumor-busting texts of the first target user ID in the target time period, Q u the comment quantity of forwarding original rumor-busting texts or publishing rumor-busting texts of the first target user ID in the target time period, Y u the like quantity of forwarding original rumor-busting texts or publishing rumor-busting texts of the first target user ID in the target time period; the user content quality coefficient of the first target user ID is determined according to the fourth formula through the user creativity coefficient and the user text quality coefficient; the fourth formula is: M u = C u × Q u ; wherein, M u is the user content quality coefficient of the first target user ID, C u is the user creativity coefficient of the first target user ID, Q u is the user text quality coefficient of the first target user ID; the influence characteristic value of the first target user ID is determined according to the user activity coefficient of the first target user ID and the user content quality coefficient through the fifth formula; the fifth formula is: wherein, A u is the influence characteristic value of the first target user ID, H u is the user activity coefficient of the first target user ID, M u is the user content quality coefficient of the first target user ID, MAXH u is the maximum user activity coefficient of the first target user ID, MAXM u is the maximum user content quality coefficient of the first target user ID; whether the influence characteristic value of the first target user ID is greater than a preset influence characteristic value is judged; if yes, the first target user ID is determined as the second target user ID. The influence characteristic value of the second target user ID is used to represent the influence of the second target user ID. Due to the openness of the social network, in the rumor-busting information transmission process, the ordinary user nodes have a certain dependence on the key user nodes or the user nodes with large transmission influence, and a cluster with the key nodes or the user nodes with large transmission influence as the center is formed. The key user nodes and the user nodes with large transmission influence have an influence and a guiding effect on the ordinary user nodes, which is similar to the relationship between "leaders" and "followers" in the group, which is the embodiment of the group power structure in the group theory. The above method screens the second target user ID with the influence greater than the preset influence.
[0198] III. The terminal screens the key rumor-buster ID in the second target user ID through the key rumor-buster identification model, comprising: determining the user event prestige of the second target user ID according to the sixth formula; the sixth formula is: wherein, I u is the user event prestige coefficient of the second target user ID, R u is the number of all original rumor-busting texts published and / or forwarded by the second target user ID in the target time period that are forwarded and / or commented on by other users, Z u is the number of all original rumor-busting texts published and / or forwarded by the second target user ID that are liked by other users, h is the target time period; the user participation coefficient of the second target user ID is determined according to a seventh formula; the seventh formula is: F u = wherein, F u is the user participation coefficient of the second target user ID, t u is the number of original rumor-busting texts commented on and / or forwarded by the second target user ID, T is the number of all original rumor-busting texts; the relationship feature value of the second target user ID is determined according to an eighth formula through the user event prestige coefficient and the user participation coefficient; the eighth formula is: wherein, D u is the relationship feature value of the second target user ID, I u is the user event prestige coefficient of the second target user ID, F u is the user participation coefficient of the second target user ID, maxI u is the maximum value of the user event prestige coefficient of the second target user ID, maxF u is the maximum value of the user participation coefficient of the second target user ID; whether the relationship feature value of the second target user ID is greater than a preset relationship feature value is determined; if yes, the second target user ID is determined as the key rumor-buster ID. The user prestige coefficient is the forwarding, commenting and liking of the comment information of other users on the second target user ID, the more the number of comments, forwards and likes of other users on the second target user ID, the higher the prestige of the second target user ID. The relationship feature value reflects the degree of cooperation of users in the social media, similar to the interpersonal relationship structure in the offline group, the analysis of user relationship data information can not only reveal the relationship network between each user node in the rumor-busting information propagation process, but also help rumor managers better master the propagation path and method of rumor-busting information, which has important significance for rumor monitoring and rumor-busting.
[0199] 210. The terminal verifies the effectiveness of the key rumor-buster identification model based on the key rumor-buster ID.
[0200] In the embodiment, the terminal verifies the effectiveness of the key debunker identification model based on the key debunker ID, which includes: selecting the identification results of the top 100 key debunker IDs as the standard, and comparing the identification results with other several representative pipe user mining algorithms PageRank, SPEAR, Leader-Rank, and UserRank in the mining in terms of recall, precision, and F1-score. The identification results of the top 100 key debunker IDs are the judgment results of the experts in the related field. Wherein, The recall is The precision is Wherein, TP represents the number of true positives in the identification results of the 100 key debunker IDs, FN represents the number of false negatives in the identification results of the 100 key debunker IDs, and FP represents the number of false positives corresponding to the identification results of the 100 key debunker IDs. The key debunker identification model has better recall, precision, and F1-score than other existing models.
[0201] The above describes a method for identifying a key debunker of social media provided by the application. Next, a device for identifying a key debunker of social media provided by the application is described.
[0202] Please refer to Figure 3 , Figure 3 An embodiment of a device for identifying a key debunker of social media provided by the application includes:
[0203] An acquisition unit 301 is configured to acquire original debunking texts and acquire user IDs that forward and / or comment on the original debunking texts.
[0204] A first identification unit 302 is configured to identify and classify text topics of the original debunking texts.
[0205] A second identification unit 303 is configured to identify a sound studio set corresponding to each type of text topic. The sound studio set includes at least two target user IDs. The target user ID is a user ID of a user participating in commenting or forwarding at least two same original debunking texts in the text topic.
[0206] A first construction unit 304 is configured to construct a user network according to the user IDs. The user network is a network established by taking the user IDs as nodes and the interaction relationships between different user IDs as edges.
[0207] The second construction unit 305 is configured to construct an event network according to the original rumor-busting text, the event network comprising an original rumor-busting text set, a derived text set of the original rumor-busting text, a text theme set of the original rumor-busting text, and a topic set corresponding to the text theme set;
[0208] The third construction unit 306 is configured to construct an echo chamber network according to the echo chamber, the echo chamber network being a network with the echo chamber as a node and a common original rumor-busting text between different echo chambers as an edge;
[0209] The fourth construction unit 307 is configured to construct a three-layer echo chamber network structure according to the user network, the event network, and the echo chamber network, the echo chamber network structure being a network structure in which the user network, the event network, and the echo chamber network are associated with each other;
[0210] The screening unit 308 is configured to construct a key rumor-buster identification model according to the echo chamber network structure, and screen a key rumor-buster ID in the echo chamber through the key rumor-buster identification model, the key rumor-buster ID being a target user ID with a sentiment feature meeting a preset sentiment feature, an influence feature value greater than a preset influence feature value, and a relationship feature value greater than a preset relationship feature value.
[0211] In the system of this embodiment, the functions performed by each unit correspond to the steps in the method embodiment described above, and thus will not be described again in detail here. Figure 1
[0212] Next, a social media key rumor-buster identification device provided by the present application will be described in detail. Figure 4 , Figure 4 Another embodiment of the social media key rumor-buster identification device provided by the present application is provided, and the identification device comprises:
[0213] The acquisition unit 401 is configured to acquire original rumor-busting texts, and acquire user IDs that forward and / or comment on the original rumor-busting texts;
[0214] The first identification unit 402 is configured to identify and classify text themes of the original rumor-busting texts.
[0215] The second identification unit 403 is configured to identify a set of echo chambers corresponding to each type of text theme, the echo chambers in the set of echo chambers comprising at least two target user IDs, the target user ID being a user ID that participates in commenting or forwarding at least two same original rumor-busting texts in the text theme.
[0216] The first construction unit 404 is configured to construct a user network according to the user IDs, the user network being a network with the user IDs as nodes and an interaction relationship between different user IDs as edges.
[0217] The second construction unit 405 is configured to construct an event network according to the original rumor-busting texts, the event network comprising a set of original rumor-busting texts, a set of derivative texts of the original rumor-busting texts, a set of text topics of the original rumor-busting texts, and a set of topics corresponding to the set of text topics;
[0218] The third construction unit 406 is configured to construct an echo chamber network according to the echo chambers, the echo chamber network being a network with the echo chambers as nodes and common original rumor-busting texts between different echo chambers as edges;
[0219] The fourth construction unit 407 is configured to construct a three-layer echo chamber network structure according to the user network, the event network, and the echo chamber network, the echo chamber network structure being a network structure in which the user network, the event network, and the echo chamber network are associated with each other;
[0220] The screening unit 408 is configured to construct a key rumor-buster identification model according to the echo chamber network structure, and screen a key rumor-buster ID in the echo chamber through the key rumor-buster identification model, the key rumor-buster ID being a target user ID with a sentiment feature meeting a preset sentiment feature, an influence feature value greater than a preset influence feature value, and a relationship feature value greater than a preset relationship feature value.
[0221] Optionally, the second identification unit 403 is specifically configured to:
[0222] obtain the at least two types of text topics classified;
[0223] fill at least two original rumor-busting texts under the first type of text topic into a first queue;
[0224] extract a first original rumor-busting text from the first queue, determine a first user ID set corresponding to the first original rumor-busting text, the first user ID set being a set of user IDs forwarding and / or commenting on the first original rumor-busting text;
[0225] respectively match the first user ID set with other user ID sets corresponding to at least one other original rumor-busting text in the first queue until the first user ID set is matched with all the other user ID sets corresponding to all the other original rumor-busting texts in the first queue, and obtain a plurality of first intersections of the first user ID set and the other user ID sets;
[0226] respectively match a second user ID set of a second original rumor-busting text with other user ID sets corresponding to at least one other original rumor-busting text in the first queue until the second user ID set is matched with all the other user ID sets corresponding to all the other original rumor-busting texts in the first queue, and obtain a plurality of second intersections of the second user ID set and the other user ID sets;
[0227] determine a plurality of first echo chamber sets corresponding to the first text subject according to the first intersection and the second intersection;
[0228] fill the original rumor-busting text under the second text subject into the second queue, and determine a plurality of second echo chamber sets corresponding to the second text subject.
[0229] Optionally, the screening unit 408 is specifically configured to:
[0230] screen a first target user ID in the target user ID in the echo chamber through the key rumor-buster identification model, and the emotional feature of the first target user ID meets a preset emotional feature;
[0231] screen a second target user ID in the first target user ID through the key rumor-buster identification model, and the influence feature of the second target user ID is greater than a preset influence feature;
[0232] screen a key rumor-buster ID in the second target user ID through the key rumor-buster identification model, and the relationship feature of the key rumor-buster ID meets a preset relationship feature.
[0233] Optionally, the screening unit 408 is specifically configured to:
[0234] obtain derivative text in which the target user ID comments on the original rumor-busting text;
[0235] analyze the emotional polarity of the derivative text corresponding to the target user ID through an emotional tendency analysis interface, and the emotional polarity includes a positive polarity and a negative polarity;
[0236] determine whether there is at least one piece of derivative text with a negative polarity in the derivative text of the target user ID;
[0237] if yes, isolate the target user ID;
[0238] if no, determine that the target user ID is the first target user ID.
[0239] Optionally, the screening unit is specifically configured to:
[0240] determine the number of echo chambers in which the first target user ID participates through the echo chamber network;
[0241] and determine the user activity coefficient of the first target user ID through a first formula according to the number of echo chambers, and the first formula is:
[0242]
[0243] wherein, H u is the user activity coefficient of the first target user ID, P is the number of all text subjects, ECS is the number of all echo chambers, and EC uThe number of echo chambers in which the first target user ID participates, p u The number of text topics in which the first target user ID participates;
[0244] The user creativity coefficient of the first target user ID is determined according to a second formula;
[0245] The second formula is:
[0246] Wherein, C u The user creativity coefficient of the first target user ID, N u The number of times that the first target user ID forwards or posts original rumor-busting texts in a target time period, h is the target time period;
[0247] The user text quality coefficient of the first target user ID is determined according to the number of times that the first target user ID forwards or posts original rumor-busting texts in a target time period and a third formula;
[0248] The third formula is:
[0249] Wherein, Q u The user text quality coefficient, N u The number of times that the first target user ID forwards or posts original rumor-busting texts in a target time period, L u The number of forwards of the first target user ID forwarding or posting original rumor-busting texts in a target time period, Q u The number of comments of the first target user ID forwarding or posting original rumor-busting texts in a target time period, Y u The number of likes of the first target user ID forwarding or posting original rumor-busting texts in a target time period;
[0250] The user content quality coefficient is determined according to a fourth formula through the user creativity coefficient and the user text quality coefficient;
[0251] The fourth formula is:
[0252] M u =C u ×Q u ;
[0253] Wherein, M u The user content quality coefficient of the first target user ID, C u The user creativity coefficient of the first target user ID, Q u The user text quality coefficient of the first target user ID;
[0254] The influence characteristic value of the first target user ID is determined according to the user activity coefficient and the user content quality coefficient of the first target user ID by a fifth formula;
[0255] The fifth formula is:
[0256]
[0257] Wherein, A u is the influence characteristic value of the first target user ID, H u is the user activity coefficient of the first target user ID, M u is the user content quality coefficient of the first target user ID, MAXH u is the maximum user activity coefficient in the first target user ID, MAXM u is the maximum user content quality coefficient in the first target user ID;
[0258] It is judged whether the influence characteristic value of the first target user ID is greater than a preset influence characteristic value;
[0259] If yes, the first target user ID is determined as the second target user ID.
[0260] Optionally, the screening unit is specifically configured to:
[0261] A user event prestige coefficient of the second target user ID is determined according to a sixth formula;
[0262] The sixth formula is:
[0263]
[0264] Wherein, I u is the user event prestige coefficient of the second target user ID, R u is the number of all original rumor-busting texts published and / or forwarded by the second target user ID in a target time period and forwarded and / or commented by other users, Z u is the number of all original rumor-busting texts published and / or forwarded by the second target user ID and liked by other users, and h is the target time period;
[0265] A user participation coefficient of the second target user ID is determined according to a seventh formula;
[0266] The seventh formula is:
[0267]
[0268] Wherein, F u is the user participation coefficient of the second target user ID, and t ua number of original debunking texts participating in forwarding and / or comments for the second target user ID, T is a number of all original debunking texts;
[0269] The relationship feature value of the second target user ID is determined according to an eighth formula through the user event prestige coefficient and the user participation coefficient.
[0270] The eighth formula is:
[0271]
[0272] Wherein, D u is the relationship feature value of the second target user ID, I u is the user event prestige coefficient of the second target user ID, F u is the user participation coefficient of the second target user ID, maxI u is the maximum value of the user event prestige coefficient of the second target user ID, maxF u is the maximum value of the user participation coefficient of the second target user ID.
[0273] It is judged whether the relationship feature value of the second target user ID is greater than a preset relationship feature value.
[0274] If yes, it is determined that the second target user ID is a key debunking user ID.
[0275] Optionally, the identification device further comprises:
[0276] The verification unit 409 is configured to verify the effectiveness of the key debunking user identification model based on the key debunking user ID.
[0277] Optionally, the identification device further comprises:
[0278] The preprocessing unit 410 is configured to preprocess the original debunking texts, and the preprocessing includes filtering duplicate and / or invalid original debunking texts.
[0279] The first identification unit 402 is specifically configured to:
[0280] Identify and classify the text topics of the preprocessed original debunking texts.
[0281] Optionally, the first identification unit 402 is specifically configured to:
[0282] Identify the text topics of the original debunking texts through the LAD topic space model, and classify the text topics, wherein the LAD topic space model has a "document-topic-word" three-layer generative Bayesian network structure.
[0283] In the system of the embodiment, the functions performed by each unit are the same as those of the above-mentioned Figure 2The steps in the method embodiments correspond, and details are not repeated here.
[0284] The application also provides a social media key debunker identification device, please refer to Figure 5 , Figure 5 An embodiment of the social media key debunker identification device provided by the application comprises:
[0285] The processor 501, the memory 502, the input and output unit 503, and the bus 504;
[0286] The processor 501 is connected with the memory 502, the input and output unit 503, and the bus 504;
[0287] The memory 502 stores a program, and the processor 501 calls the program to execute any one of the above social media key debunker identification methods.
[0288] The application also relates to a computer readable storage medium, and the computer readable storage medium stores a program, when the program runs on the computer, so that the computer executes any one of the above social media key debunker identification methods.
[0289] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device and unit can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0290] In the several embodiments provided by the application, it should be understood that the disclosed system, device and method can be implemented by other ways. For example, the device embodiments described above are only schematic, and for example, the division of the units is only a logical function division, and there can be another division way in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection between the units can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0291] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment scheme.
[0292] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.
[0293] When the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, read-only memory), a random access memory (RAM, random access memory), a magnetic disk or an optical disk, and various media that can store program codes.
Claims
1. A method for identifying a social media key debunker, characterized in that, The identification method comprises: obtaining original refutation texts, and obtaining user IDs forwarding and / or commenting on the original refutation texts; identifying and classifying text topics of the original refutation texts; obtaining at least two types of classified text topics; filling at least two original refutation texts under a first type of text topic into a first queue; extracting a first original refutation text from the first queue, determining a first user ID set corresponding to the first original refutation text, the first user ID set being a user ID set forwarding and / or commenting on the first original refutation text; respectively matching the first user ID set with other user ID sets corresponding to at least one other original refutation text in the first queue until the first user ID set is matched with all other user ID sets corresponding to all other original refutation texts in the first queue, obtaining a plurality of first intersections of the first user ID set and the other user ID sets; respectively matching a second user ID set of a second original refutation text with other user ID sets corresponding to at least one other original refutation text in the first queue until the second user ID set is matched with all other user ID sets corresponding to all other original refutation texts in the first queue, obtaining a plurality of second intersections of the second user ID set and the other user ID sets; determining a plurality of first echo chamber sets corresponding to the first type of text topic according to the first intersections and the second intersections; filling original refutation texts under a second type of text topic into a second queue, and determining a plurality of second echo chamber sets corresponding to the second type of text topic; constructing a user network according to the user IDs, the user network being a network with the user IDs as nodes and interaction relationships between different user IDs as edges; constructing an event network according to the original refutation texts, the event network comprising a set of original refutation texts, a set of derivative texts of the original refutation texts, a set of text topics of the original refutation texts, and a set of topics corresponding to the set of text topics; constructing an echo chamber network according to the echo chambers, the echo chamber network being a network with the echo chambers as nodes and common original refutation texts between different echo chambers as edges; constructing a three-layer echo chamber network structure according to the user network, the event network, and the echo chamber network, the echo chamber network structure being a network structure in which the user network, the event network, and the echo chamber network are associated with each other; constructing a key refuter identification model according to the echo chamber network structure, and screening key refuter IDs in the echo chambers through the key refuter identification model, the key refuter ID being a target user ID with a sentiment feature meeting a preset sentiment feature, an influence feature value greater than a preset influence feature value, and a relationship feature value greater than a preset relationship feature value.
2. The identification method according to claim 1, characterized in that, The screening of the key refuter IDs in the echo chambers through the key refuter identification model comprises: screening a first target user ID in the target user ID in the echo chamber through the key debunking person identification model, the emotional feature of the first target user ID meeting a preset emotional feature; screening a second target user ID in the first target user ID through the key debunking person identification model, the influence feature of the second target user ID being greater than a preset influence feature; screening a key debunking person ID in the second target user ID through the key debunking person identification model, the relationship feature of the key debunking person ID meeting a preset relationship feature.
3. The identification method according to claim 2, characterized in that, The screening of the first target user ID in the target user ID in the echo chamber through the key debunking person identification model includes: obtaining a derivative text of the target user ID commenting on the original debunking text; analyzing the emotional polarity of the derivative text corresponding to the target user ID through an emotional tendency analysis interface, the emotional polarity including a positive polarity and a negative polarity; determining whether there is at least one piece of derivative text with a negative polarity in the derivative text of the target user ID; if yes, isolating the target user ID; if no, determining the target user ID as the first target user ID.
4. The identification method according to claim 3, characterized in that, The screening of the second target user ID in the first target user ID through the key debunking person identification model includes: screening a second target user ID in the first target user ID through the key debunking person identification model, the influence feature value of the second target user ID being greater than a preset influence feature value; identifying the number of echo chambers participated by the first target user ID through the echo chamber network; wherein H u is the user activity coefficient of the first target user ID, P is the number of all text topics, ECS is the number of all echo chambers, EC u is the number of echo chambers in which the first target user ID participates, p u is the number of text topics in which the first target user ID participates; and determining the user activity coefficient of the first target user ID through a first formula according to the number of echo chambers, the first formula being: The second formula is: Wherein, C u The creativity coefficient of the user with the first target user ID, N u The number of times that the first target user ID forwards or publishes the original rumor-busting text within a target time period, h is the target time period. determining the user creativity coefficient of the first target user ID according to a second formula; The third formula is: wherein Q u is a user text quality coefficient, N u is the number of times the first target user ID forwards original rumor-busting texts or publishes rumor-busting texts within a target time period, L u is the forwarding quantity of the first target user ID forwarding original rumor-busting texts or publishing rumor-busting texts within a target time period, Q u is the comment quantity of the first target user ID forwarding original rumor-busting texts or publishing rumor-busting texts within a target time period, Y u is the like quantity of the first target user ID forwarding original rumor-busting texts or publishing rumor-busting texts within a target time period. determining the user text quality coefficient of the first target user ID according to the number of original debunking texts forwarded or published by the first target user ID within a target time period and a third formula; determining the user content quality coefficient through a fourth formula according to the user creativity coefficient and the user text quality coefficient; M u = C u x Q u ; wherein M u is a user content quality coefficient for the first target user ID, C u is a user creativity coefficient for the first target user ID, Q u is a user text quality coefficient for the first target user ID. the fourth formula being: determining the influence feature value of the first target user ID through a fifth formula according to the user activity coefficient and the user content quality coefficient of the first target user ID; wherein A u is an influence feature value of the first target user ID, H u is a user activity coefficient of the first target user ID, M u is a user content quality coefficient of the first target user ID, MAXH u is the largest user activity coefficient in the first target user ID, MAXM u is the largest user content quality coefficient in the first target user ID; the fifth formula being: determining whether the influence feature value of the first target user ID is greater than the preset influence feature value; 5. The identification method according to claim 4, characterized in that, if yes, determining the first target user ID as the second target user ID. The screening of the key debunking person ID in the second target user ID through the key debunking person identification model includes: determining the user event prestige coefficient of the second target user ID according to a sixth formula; wherein, I u is the user event influence coefficient of the second target user ID, R u is the number of all original rumor-busting texts published and / or forwarded by the second target user ID in the target time period being forwarded and / or commented on by other users, Z u is the number of all original rumor-busting texts published and / or forwarded by the second target user ID being liked by other users, h is the target time period; the sixth formula being: determining the user participation degree coefficient of the second target user ID according to a seventh formula; wherein F u is the user engagement coefficient of the second target user ID, t u is the number of original debunking texts that the second target user ID participated in forwarding and / or commenting, T is the number of all original debunking texts; the seventh formula being: determining the relationship feature value of the second target user ID through an eighth formula according to the user event prestige coefficient and the user participation degree coefficient; The eighth formula is: wherein D u is a relationship feature value of the second target user ID, I u is a user event influence coefficient of the second target user ID, F u is a user engagement coefficient of the second target user ID, maxI u is a maximum value of the user event influence coefficient of the second target user ID, maxF u is a maximum value of the user engagement coefficient of the second target user ID. determining whether the relationship feature value of the second target user ID is greater than a preset relationship feature value; if yes, determining that the second target user ID is a key rumor-busting ID.
6. The identification method according to any one of claims 1 to 5, characterized in that, After the key rumor-busting ID recognition model is constructed according to the echo chamber network structure and the key rumor-busting ID of the echo chamber is screened through the key rumor-busting ID recognition model, the recognition method further comprises: verifying the effectiveness of the key rumor-busting ID recognition model based on the key rumor-busting ID.
7. The identification method according to any one of claims 1 to 5, characterized in that, Before the text theme of the original rumor-busting text is recognized and the text theme is classified, the recognition method further comprises: preprocessing the original rumor-busting text, the preprocessing comprising filtering duplicate and / or invalid original rumor-busting texts; the recognition and classification of the text theme of the original rumor-busting text comprises: recognizing and classifying the text theme of the preprocessed original rumor-busting text.
8. The identification method of claim 1, wherein, The recognition and classification of the text theme of the original rumor-busting text comprises: recognizing the text theme of the original rumor-busting text and classifying the text theme through a LAD theme space model, the LAD theme space model having a "document-theme-word" three-layer generative Bayesian network structure.
9. A device for identifying key debunkers on social media, characterized in that, The recognition device comprises: an acquisition unit configured to acquire original rumor-busting texts and user IDs forwarding and / or commenting on the original rumor-busting texts; a first recognition unit configured to recognize and classify the text theme of the original rumor-busting text; a second recognition unit configured to acquire at least two types of classified text themes; fill at least two original rumor-busting texts under a first type of text theme into a first queue; extract a first original rumor-busting text from the first queue, determine a first user ID set corresponding to the first original rumor-busting text, the first user ID set being a user ID set forwarding and / or commenting on the first original rumor-busting text; match the first user ID set with other user ID sets corresponding to at least one other original rumor-busting text in the first queue respectively until the first user ID set is matched with all other user ID sets corresponding to all other original rumor-busting texts in the first queue, acquire a plurality of first intersections of the first user ID set and the other user ID sets; match a second user ID set of a second original rumor-busting text with other user ID sets corresponding to at least one other original rumor-busting text in the first queue respectively until the second user ID set is matched with all other user ID sets corresponding to all other original rumor-busting texts in the first queue, acquire a plurality of second intersections of the second user ID set and the other user ID sets; determine a plurality of first echo chamber sets corresponding to the first type of text theme according to the first intersections and the second intersections; fill original rumor-busting texts under a second type of text theme into a second queue and determine a plurality of second echo chamber sets corresponding to the second type of text theme; A first constructing unit is configured to construct a user network according to the user IDs, the user network being a network with the user IDs as nodes and interaction relationships between different user IDs as edges; A second constructing unit is configured to construct an event network according to the original rumor-busting texts, the event network including a set of original rumor-busting texts, a set of derivative texts of the original rumor-busting texts, a set of text topics of the original rumor-busting texts, and a set of topics corresponding to the set of text topics; A third constructing unit is configured to construct an echo chamber network according to the echo chambers, the echo chamber network being a network with the echo chambers as nodes and common original rumor-busting texts between different echo chambers as edges; A fourth constructing unit is configured to construct a three-layer echo chamber network structure according to the user network, the event network, and the echo chamber network, the echo chamber network structure being a network structure in which the user network, the event network, and the echo chamber network are associated with each other; A screening unit is configured to construct a key rumor-buster identification model according to the echo chamber network structure, and to screen key rumor-buster IDs in the echo chambers through the key rumor-buster identification model, the key rumor-buster IDs being target user IDs with emotional features meeting preset emotional features, influence feature values greater than a preset influence feature value, and relationship feature values greater than a preset relationship feature value.
Citation Information
Patent Citations
Network rumor propagation control method based on representation learning
CN110795641A
Rumor propagation control method based on rumor promotion-rumor refuting message and representation learning
CN110825948A