Inference device and inference method
The estimation device and method address the limitations of conventional techniques by quantitatively evaluating misinformation spread and user intervention effectiveness within social networks, enabling targeted interventions to reduce misinformation by identifying key users.
Patent Information
- Application Number
- PCT/JP2025/008001
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-07
- Filing Date
- 2025-03-05
- Publication Date
- 2025-12-11
AI Technical Summary
Conventional techniques for evaluating the effectiveness of user intervention strategies against misinformation on social media fail to account for individual user influences and social relationships, leading to ineffective intervention strategies at the group level.
An estimation device and method that estimates misinformation sharing behavior of each user using social network data, constructs a social network of friendships, and evaluates the intervention effect within this network to identify key users for targeted interventions.
Enables quantitative evaluation of misinformation spread and effective user intervention strategies by identifying users who can significantly reduce misinformation propagation, allowing for optimized group-level interventions.
Smart Images

Figure JP2025008001_11122025_PF_FP_ABST
Abstract
Description
Estimation device and estimation method
[0001] The present invention relates to an estimation device and an estimation method.
[0002] Social media is widely used by many people as a platform for online information dissemination and communication. However, in recent years, the spread of misinformation on social media has had serious negative effects on a wide range of areas, from politics to public health, including hindering the formation of healthy public opinion, distorting election results, and encouraging ineffective or harmful self-medication practices.
[0003] As a way to address the problem of misinformation on social media, various user intervention techniques are being researched, such as presenting fact-checking articles, media literacy education, and nudges. These user intervention techniques are known to have a certain level of intervention effectiveness at the individual level, such as reducing users' cognitive vulnerability to misinformation (ease with which they are gullible).
[0004] However, whether an individual shares misinformation is generally strongly influenced not only by their own cognitive factors but also by factors such as social relationships between individuals. Therefore, to effectively counter the spread of misinformation, it is necessary to quantitatively evaluate the intervention effect at the group level (e.g., how much the intervention reduces the spread of misinformation within a group) that takes into account social relationships between individuals, rather than just the individual level (e.g., how much the intervention reduces the spread of misinformation within a group), and then design an optimal intervention strategy.
[0005] A technique for evaluating the effectiveness of user intervention techniques at a group level has been proposed (Non-Patent Document 1). In the conventional technique, first, the amount of posts about a certain misinformation content on social media at time t is calculated as y t We build a statistical model to predict y and fit various parameters of the statistical model based on actual social media data. Next, we use numerical simulations to evaluate whether user intervention techniques such as fact-checking, nudges, and account suspensions can affect the amount of posts related to misinformation content. tEvaluate how much it reduces
[0006] Joseph B. Bak-Coleman, Ian Kennedy, Morgan Wack, et al., "Combining interventions to reduce the spread of viral misinformation," Nat Hum Behav 6, 1372-1380 (2022).
[0007] However, conventional techniques directly model the volume of misinformation posts across social media, ignoring factors such as the influence of individual users, social relationships, and individual differences in intervention effectiveness. As a result, conventional techniques cannot capture the impact of interventions by micro-users on macro-groups, i.e., the impact of interventions on an individual on the entire group. Therefore, conventional techniques have the drawback of being unable to help design intervention strategies to combat the spread of misinformation.
[0008] The present invention has been made in consideration of the above, and aims to provide an estimation device and an estimation method that can evaluate the spread of misinformation in a group when a certain user is involved.
[0009] In order to solve the above-mentioned problems and achieve the objectives, the estimation device of the present invention is characterized by having a first estimation unit that estimates the misinformation sharing behavior of each user using user data obtained based on social network data, and a second estimation unit that builds a social network of friendships from the user's friend data and estimates the intervention effect in the social network based on the misinformation sharing behavior of each user.
[0010] Furthermore, the estimation method of the present invention is an estimation method executed by an estimation device, and is characterized by including the steps of: estimating the misinformation sharing behavior of each user using user data obtained based on social network data; and constructing a social network of friendships from the user's friend data, and estimating the intervention effect in the social network based on the misinformation sharing behavior of each user.
[0011] According to the present invention, it is possible to evaluate the spread of misinformation in a group when a certain user is involved.
[0012] FIG. 1 is a diagram schematically illustrating an example of the configuration of an estimation device according to an embodiment. FIG. 2 is a diagram illustrating an example of a method for estimating misinformation sharing behavior. FIG. 3-1 is a diagram illustrating an example of output information of the estimation device illustrated in FIG. 1. FIG. 3-2 is a diagram illustrating an example of output information of the estimation device illustrated in FIG. 1. FIG. 4 is a flowchart illustrating the processing steps of estimation processing according to an embodiment. FIG. 5 is a diagram illustrating an example of a computer on which the estimation device is implemented by executing a program. FIG. 6 is a diagram schematically illustrating an example of the configuration of an estimation device according to a second embodiment. FIG. 7 is a flowchart illustrating the processing steps of estimation processing according to the second embodiment. FIG. 8-1 is a diagram illustrating an example of an intervention effect. FIG. 8-2 is a diagram illustrating an example of an intervention effect. FIG. 8-3 is a diagram illustrating an example of an intervention effect. FIG. 9 is a diagram schematically illustrating an example of the configuration of an estimation system according to an embodiment. FIG. 10-1 is a flowchart illustrating the processing steps of data preprocessing according to an embodiment. FIG. 10-2 is a flowchart illustrating the processing steps of data learning processing according to an embodiment.
[0013] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to this embodiment. In addition, in the description of the drawings, the same parts are designated by the same reference numerals.
[0014] In this embodiment, the social network between users is taken into consideration to estimate how and / or how much misinformation will spread in a social network when a user is involved. A social network is a social network built on the web via social media such as social networking services, electronic bulletin boards, posting sites, video distribution sites, and information sharing sites.
[0015] In the embodiment, when intervention is performed to suppress the spread of misinformation that spreads from user to user on social media, for example, the above estimation can be performed to select users who are effective in suppressing the spread of misinformation. In the embodiment, the intervention effect of the user intervention technique is quantitatively evaluated at the group level, making it possible to select intervening users who will efficiently suppress the spread of misinformation.
[0016] <Estimation Device> An estimation device according to an embodiment will be described below. Fig. 1 is a diagram schematically illustrating an example of the configuration of an estimation device according to an embodiment.
[0017] The estimation device 10 according to the embodiment is realized, for example, by loading a predetermined program into a computer or the like including a ROM (Read Only Memory), a RAM (Random Access Memory), a CPU (Central Processing Unit), etc., and causing the CPU to execute the predetermined program. The estimation device 10 also has a communication interface for transmitting and receiving various information to and from other devices connected via a network, etc. The estimation device 10 is realized by a general-purpose computer such as a workstation or a personal computer.
[0018] The estimation device 10 according to the embodiment includes a data collection unit 11 , a data processing unit 12 , and an estimation processing unit 13 .
[0019] The data collection unit 11 collects social network data, for example, from social media, by communicating with an external device. The social network data is data related to information posted, distributed, or shared on social media, and includes the content posted, distributed, or shared, user information of the source of the posting, distribution, or sharing, information of the destination of the posting, distribution, or sharing, and the conditions for posting, distribution, or sharing. The data collection unit 11 collects data by using an API (Application Programming Interface) or scraping from social media, by receiving the data from a company operating the social media, or by purchasing the data from the company operating the social media. The data collection unit 11 may collect various information related to users from users.
[0020] The data processing unit 12 acquires user data based on, for example, data collected by the data collection unit 11 (including social network data), converts it into a format (e.g., vector) that can be processed by the estimation unit 14 (described later), and then outputs it to the estimation unit 14.
[0021] Here, the user is a person who is the subject of estimation in the estimation device 10. Alternatively, the user is a person who belongs to the same group as the subject of estimation. The estimation device 10 estimates what kind of misinformation sharing behavior the user will exhibit when exposed to misinformation, and how the misinformation sharing behavior will change when receiving intervention.
[0022] User data is data about a user, and includes, for example, user activity data, user characteristic data, and friend data.
[0023] The user activity data is data indicating the user's past activities. For example, the user's postings and sharing on social media are examples of the user's activity data. The user characteristic data is data indicating the user's interests and political leanings. The friend data is data indicating the user's friendships, and in social media, users who are followed are considered to be friends. The data processing unit 12 classifies the data collected by the data collection unit 11 into, for example, user activity data, user characteristic data, friend data, and social network data.
[0024] The estimation processing unit 13 includes an estimation unit 14 and an estimation result output unit 15 that outputs the estimation result obtained by the estimation unit 14 .
[0025] The estimation unit 14 includes a user activity data unit 141 that stores user activity data, a user characteristic data unit 142 that stores user characteristic data, a social network data unit 143 that stores social network data, a user misinformation sharing behavior estimation unit 144 (first estimation unit), and a network (NW) intervention effect estimation unit 145 (second estimation unit).
[0026] The user misinformation sharing behavior estimation unit 144 estimates the misinformation sharing behavior of a user using user data acquired based on social network data. The user misinformation sharing behavior estimation unit 144 uses a machine learning model to estimate, based on user activity data and / or user characteristic data, what kind of misinformation sharing behavior a user will exhibit when exposed to misinformation and / or how the misinformation sharing behavior will change when receiving intervention.
[0027] The machine learning model used by the user misinformation sharing behavior estimation unit 144 learns, for each user, the user's past activities, the user's characteristics, and the user's misinformation sharing behavior in response to past input misinformation. The machine learning model used by the user misinformation sharing behavior estimation unit 144 receives as input the user to be estimated and the misinformation given to this user, and outputs the user's susceptibility to being deceived in this case.
[0028] Various methods can be applied as a method for the user misinformation sharing behavior estimation unit 144 to estimate misinformation sharing behavior.
[0029] 2 is a diagram illustrating an example of a method for estimating misinformation sharing behavior. For example, the user misinformation sharing behavior estimation unit 144 receives vectorized user data 31 as input and uses machine learning techniques such as a neural network 1441 to learn and build a prediction model that outputs prediction results 32, such as the probability that a user will be deceived by misinformation or the probability that a user will share misinformation (hereinafter simply referred to as vulnerability to misinformation) when a certain intervention is received / not received.
[0030] Another method for estimating misinformation sharing behavior by the user misinformation sharing behavior estimation unit 144 is to conduct a user survey of users with various attributes, identify the main factors that influence users' vulnerability to misinformation, and build a prediction model in which the influencing factors are used as explanatory variables and vulnerability to misinformation is used as a target variable. Major influencing factors include, for example, media literacy ability, political partisanship, nationality, and age.
[0031] The NW intervention effect estimation unit 145 estimates the intervention effect in the social network.
[0032] Generally, it is not economically realistic to intervene on all social media users. Therefore, when intervening on social media, only a small portion of users can be intervened. Therefore, in order to take efficient measures, it is desirable to intervene on users who can most effectively suppress the spread of misinformation on social networks.
[0033] Therefore, the NW intervention effect estimation unit 145 constructs a social network of friendships from the user's friend data. The NW intervention effect estimation unit 145 estimates the intervention effect in the social network based on the user's misinformation sharing behavior in the constructed social network.
[0034] The NW intervention effect estimation unit 145 estimates, based on the misinformation sharing behavior of each user, how much misinformation can be reduced and / or how the prevalence of misinformation will change in the constructed social network when intervention is made on a specific user or user group, depending on the specific user or user group that made the intervention. For example, the specific user or user group that made the intervention is specified by, for example, a countermeasure implementer. The prevalence of misinformation when intervention is made on a specific user (or user group) is the proportion of users exposed to misinformation in the social network.
[0035] When intervention is performed, the rate at which misinformation spreads changes depending on the user or user group that receives the intervention, making them less susceptible to misinformation. Therefore, the NW intervention effect estimation unit 145 estimates the intervention effect, which indicates how much misinformation can be reduced, based on users' misinformation sharing behavior in the constructed social network.
[0036] Various methods can be applied as a method for estimating the intervention effect in a social network by the NW intervention effect estimation unit 145. For example, the NW intervention effect estimation unit 145 estimates the intervention effect in a social network based on the misinformation sharing behavior of users by, for example, performing a simulation.
[0037] Furthermore, the NW intervention effect estimation unit 145 may perform estimation using an information diffusion model, assuming that misinformation spreads virally among users of a social network according to any information diffusion model (machine learning model). The information diffusion model is a model that estimates the diffusion of information among users of a social network. The NW intervention effect estimation unit 145 may use a method of estimating the prevalence rate of misinformation through theoretical analysis using the information diffusion model or numerical simulation.
[0038] An information diffusion model is a model that defines the mechanism by which information is transmitted from one user to another in a social network. There are various information diffusion models, such as the independent cascade model, linear threshold model, epidemic model, and point process model, but we will not limit ourselves to a specific model here.
[0039] In addition, the network intervention effect estimation unit 145 may use a method of constructing a model that predicts the prevalence rate of misinformation using machine learning techniques such as graph neural networks based on data regarding users' posting and / or sharing history of misinformation content.
[0040] Furthermore, the NW intervention effect estimation unit 145 may estimate users in the social network for whom intervention is highly effective, i.e., users whose intervention can significantly contribute to reducing the prevalence of misinformation throughout the social network, and output the estimate to a countermeasure provider. In this case, the NW intervention effect estimation unit 145 estimates users in a predetermined ranking, starting from the top, as users for whom intervention is highly effective, and outputs the estimate to a countermeasure provider. For example, if the estimation result by the user misinformation sharing behavior estimation unit 144 determines that users with a certain characteristic are relatively easily deceived, the NW intervention effect estimation unit 145 may output a user group in the social network that has this specific characteristic. For example, if the number n of users with this specific characteristic is less than a predetermined threshold k, the number of users for whom intervention is highly effective is set to n users, and if the number is equal to or greater than the threshold, the number of users for whom intervention is highly effective is set to the top k users. According to this setting, the NW intervention effect estimation unit 145 estimates the n users or the top k users for whom intervention is highly effective, and outputs the estimate to a countermeasure provider as a user group for whom intervention is highly effective.
[0041] Through the process described above, the estimation device 10 can quantitatively evaluate the effectiveness of suppressing the spread of misinformation when intervention is performed on a specific user (or user group).
[0042] The estimation result output unit 15 processes the estimation results in a format that supports intervention decisions by countermeasure actors (social media, news media, specialized institutions, etc.) who take measures against the spread of misinformation, and outputs the processed results to the countermeasure actors. The estimation device 10 processes the estimation results in a format such as a list of users who contribute significantly to the intervention effect, or a plot of a simulation result of the intervention effect when intervention is performed on a specified user. Using the processed estimation results in this way can help countermeasure actors design user intervention strategies for efficient countermeasures against the spread of misinformation.
[0043] 3A and 3B are diagrams showing examples of output information of the estimation device 10 shown in FIG.
[0044] Figure 3-1 shows an example of a list of the calculation results for each user's contribution to the intervention effect. In the example of Figure 3-1, user "@GEQpT2994" shows the largest contribution to the intervention effect. Therefore, countermeasures can determine that intervening with user "@GEQpT2994" is likely to effectively curb the spread of misinformation.
[0045] Figure 3-2 shows an example of the simulation results of the intervention effect. Figure 3-2 shows the simulation results of the prevalence of misinformation when users "@DaYVS9617" and "@NaEyL4036" are specified as intervention targets (Input) in a group of 10 people.
[0046] In the output example shown in FIG. 3-2, icons 411 and 412 represent users "@DaYVS9617" and "@NaEyL4036" who intervened. Icons 421, 422, and 423 represent users who shared false information. In this example, three out of ten users shared false information, so the prevalence rate is σ=0.3.
[0047] In this way, by using the estimation results of the estimation device 10, countermeasure personnel can determine what kind of intervention they should take and to whom in order to effectively suppress the spread of false information.
[0048] <Estimation Process> FIG. 4 is a flowchart showing the processing procedure of the estimation process according to the embodiment.
[0049] The data collection unit 11 collects various data including social network data from social media etc. (step S1). The data processing unit 12 processes the data (including social network data) collected by the data collection unit 11 (step S2), for example, and acquires user data including user activity data and user characteristics.
[0050] The user misinformation sharing behavior estimation unit 144 estimates the user's misinformation sharing behavior using user data acquired from the social network data (step S3). Subsequently, the NW intervention effect estimation unit 145 constructs a social network of friendships from the user's friend data and estimates the intervention effect in the social network based on the user's misinformation sharing behavior estimated in step S3 (step S4).
[0051] Then, the estimation result output unit 15 processes the estimation result into a form that will support the countermeasure person's decision to intervene, and outputs the processed result to the countermeasure person (step S5).
[0052] <Effects of the embodiment> Conventional user intervention techniques for countering misinformation, such as nudges and media literacy education, have only established methods for measuring the effectiveness of intervention at the user level, such as how less susceptible users are to misinformation.
[0053] In contrast, the estimation device 10 according to the embodiment estimates a user's misinformation sharing behavior using user data acquired from social network data. The estimation device 10 then constructs a social network of friendships from the user's friend data and estimates the intervention effect in the social network based on the user's misinformation sharing behavior. In other words, the estimation device 10 performs two-stage estimation: estimating each user's misinformation sharing behavior and estimating the intervention effect in the social network.
[0054] As a result, the estimation device 10 can quantitatively evaluate the intervention effect of user intervention techniques at the group level. That is, the estimation device 10 can quantitatively evaluate the spread of misinformation when an intervention is made for a certain user in a group. In other words, the estimation device can quantitatively estimate how much the spread of misinformation will change depending on who is intervened and what kind of intervention is made to whom in a group such as a social media community or network.
[0055] By using the estimation results from this estimation device 10, countermeasure providers such as social media, news media, and specialized institutions that take measures to prevent the spread of misinformation can select intervening users to efficiently suppress the spread of misinformation.
[0056] Therefore, by explicitly considering the social networks of users, the estimation device 10 can evaluate how and / or to what extent misinformation will spread if intervention is performed on a certain user. This makes it possible to design an efficient user intervention strategy, such as selecting the optimal user to intervene in order to suppress the spread of misinformation.
[0057] <System Configuration of the First and Second Embodiments> The estimation device 10 and the estimation device 20 of the second embodiment are functional concepts and do not necessarily need to be physically configured as shown in the drawings. In other words, the specific forms of distribution and integration of the functions of the estimation devices 10 and 20 are not limited to those shown in the drawings, and all or part of them can be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc.
[0058] Furthermore, all or any part of the processes performed in the estimation devices 10 and 20 may be realized by a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and a program analyzed and executed by the CPU and the GPU (Graphics Processing Unit). Furthermore, each process performed in the estimation devices 10 and 20 may be realized as hardware using wired logic.
[0059] Furthermore, among the processes described in the first and second embodiments, all or part of the processes described as being performed automatically can be performed manually. Alternatively, all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the process procedures, control procedures, specific names, and information including various data and parameters described above and illustrated can be changed as appropriate unless otherwise specified.
[0060] 5 is a diagram showing an example of a computer in which the estimation devices 10 and 20 are realized by executing a program. The computer 1000 has, for example, a memory 1010 and a CPU 1020. The computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0061] The memory 1010 includes a ROM 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.
[0062] The hard disk drive 1090 stores, for example, an operating system (OS) 1091, an application program 1092, a program module 1093, and program data 1094. That is, the programs that define the processes of the estimation devices 10 and 20 are implemented as program modules 1093 in which code executable by the computer 1000 is written. The program modules 1093 are stored, for example, in the hard disk drive 1090. For example, the program modules 1093 for executing processes similar to the functional configurations of the estimation devices 10 and 20 are stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced with an SSD (Solid State Drive).
[0063] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in memory 1010 or hard disk drive 1090. Then, CPU 1020 reads out program module 1093 or program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as necessary and executes them.
[0064] The program module 1093 and program data 1094 may not necessarily be stored in the hard disk drive 1090, but may also be stored in a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or a wide area network (WAN)). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.
[0065] <Second Embodiment> In the second embodiment, when intervening to suppress the spread of misinformation that spreads from user to user on social media, for example, the estimation described below can be performed to select users who are effective in suppressing the spread of misinformation. Note that misinformation naturally includes content such as unverified rumors, information from unreliable sources, conspiracy theories, satire, and hate speech.
[0066] There are various interventions to address the problem of misinformation on social media, but pre-debunking is seen as promising. Debunking is an intervention that aims to correct users' erroneous beliefs by correcting the errors in the misinformation or presenting fact-checking information after users have been exposed to misinformation. Examples of such interventions include presenting fact-check articles, displaying warning labels, and presenting community notes.
[0067] Prebanking is an intervention that encourages users to acquire psychological resistance to misinformation by presenting warnings, counterarguments, and tactics of misinformation before (or immediately before or after) exposure to misinformation. Examples of such interventions include literacy education, inoculation, and nudges.
[0068] The individual effects of prebanking on each user and the optimal prebanking method may differ depending on the characteristics of the user and the content. For example, literacy education is likely to be ineffective for users who already have high literacy levels. Also, for example, intervention is unlikely to work for users with strong sharing motivations (such as financial or political motivations).
[0069] In contrast, the second embodiment provides an estimation device and an estimation method for efficiently selecting users who are targets for intervention in order to maximize the effectiveness of pre-banking in order to minimize the spread of misinformation on social media.
[0070] <Estimation Device> An estimation device according to a second embodiment will be described below. Fig. 6 is a diagram schematically illustrating an example of the configuration of the estimation device according to the second embodiment.
[0071] The estimation device 20 according to the second embodiment is realized by, for example, loading a predetermined program into a computer or the like including a ROM (Read Only Memory), a RAM (Random Access Memory), a CPU (Central Processing Unit), etc., and causing the CPU to execute the predetermined program. The estimation device 20 also has a communication interface for transmitting and receiving various types of information to and from other devices connected via a network, etc. The estimation device 20 is realized by a general-purpose computer such as a workstation or a personal computer.
[0072] The estimation device 20 according to the second embodiment includes a data collection unit 21, a social media DB 22, and an estimation processing unit 23. The estimation processing unit 23 includes an estimation unit 24, a diffusion model generation unit 25, a target selection unit 26, and a result output unit 27.
[0073] The data collection unit 21 communicates with an external device to collect, for example, social media data, which is data related to information posted, distributed, and shared on social media, from social media. The data collection unit 21 may also collect user profiles, past posts and shared information, follow information, and the like from social media. The data collection unit 21 stores the collected data in the social media DB 22 as social media data. The data collection unit 21 outputs the social media data stored in the social media DB 22 to the estimation unit 24. The social media DB 22 may aggregate the user activity data 141, user characteristic data 142, and social network data 143 shown in FIG. 1 and store the aggregated data as social media data.
[0074] Here, the user is a person to be estimated by the estimation device 20, and is a person who can become an intervention target user (a user subject to intervention). The estimation unit 24 estimates each user's susceptibility to misinformation obtained based on collected data when each user is exposed to misinformation, the propagation probability of the misinformation spreading from each user to other users, and the individual intervention effect for each user in the social network. The estimation unit 24 includes an individual intervention effect estimation unit 241, a user influence estimation unit 242, and a propagation probability estimation unit 243.
[0075] The individual intervention effect estimation unit 241 calculates an individual intervention effect for a certain user based on the user's social media data. For example, the individual intervention effect estimation unit 241 estimates the individual intervention effect ε(u, X) when intervention X is performed on user u based on the social media data. The individual intervention effect ε(u, X) indicates, for example, how less likely user u will be deceived by misinformation as a result of intervention X.
[0076] In addition to social media data, the individual intervention effect estimation unit 241 may collect, for example, user survey data 33 for identifying user factors that are the main factors behind the effect of intervention X. The user survey data 33 may be, for example, data obtained by conducting a randomized controlled trial using a crowdsourcing service or the like to investigate the correlation between intervention X and user factors, targeting users who currently use or have previously used social media. The user survey data 33 may also be public data investigating the relationship between intervention X and user factors.
[0077] The individual intervention effect estimation unit 241 analyzes the correlation between the effect size of intervention X and user factors by statistical analysis such as stratified analysis and causal inference based on the social media data and / or user survey data 33. For example, in stratified analysis, the individual intervention effect estimation unit 241 may divide a certain user factor into several strata (e.g., high / middle / low literacy), calculate the difference in the average outcome between the intervention group and the non-intervention group in each strata as the intervention effect, and calculate the intervention effect in each strata.
[0078] The individual intervention effect estimation unit 241 analyzes the correlation between the effect amount of intervention X and user factors, and then estimates the user factors whose correlations were identified in the previous step from the user's social media data using machine learning or the like, and predicts the individual intervention effect ε(u, X) of intervention X for each user. For example, take a case where the results of analyzing the correlations of user factors in the previous step indicate that the effect of intervention X is correlated with user factor A (correlation coefficient r). In this case, user factor A(u) of user u may be estimated from social media data using text analysis, machine learning, or the like, and the individual intervention effect ε(u, X) may be calculated using the following formula (P): ε(u, X) = r × A(u) (P)
[0079] The user influence estimation unit 242 calculates the user's susceptibility to misinformation based on the user's social media data. For example, the user influence estimation unit 242 calculates the user influence s(u, c) of user u on misinformation c. The user influence s(u, c) indicates the user u's sensitivity and is an index of the user u's susceptibility to misinformation c. The higher the user influence s(u, c), the more likely the user u is to be fooled by misinformation c and share the misinformation c. The lower the user influence s(u, c), the less likely the user u is to be fooled by misinformation c and the more likely the user u is to share a rebuttal to the misinformation c.
[0080] The user influence estimation unit 242 may calculate the proportion of false information among the contents previously shared by the user u, and calculate the user influence s(u, c). Alternatively, the user influence estimation unit 242 may construct a logistic regression model or a machine learning model that predicts the probability that the user u will share false information from the social media DB 22, and calculate the user influence s(u, c) of the user as the predicted value.
[0081] The propagation probability estimation unit 243 calculates the probability that misinformation shared by a user will propagate to another user based on past misinformation diffusion data. For example, the propagation probability estimation unit 243 calculates the propagation probability p(u, v, c) of misinformation c shared by a user u to another user v based on the user's characteristics and past misinformation c diffusion data.
[0082] The propagation probability estimation unit 243 may calculate the propagation probability p(u, v, c) using a simple frequency (the proportion of misinformation c viewed and shared by user v among the misinformation c shared by user u). The propagation probability estimation unit 243 may estimate the propagation probability p(u, v, c) using maximum likelihood estimation, Bayesian estimation, a machine learning model, or the like based on the user's characteristics and past diffusion data of misinformation c. For example, the user's characteristics include the user's social media usage frequency, sharing frequency, personality, interest in misinformation c, etc.
[0083] The diffusion model generation unit 25 uses the estimation result by the estimation unit 24 to generate a network diffusion model M that estimates the propagation probability p of false information c from one user to another when intervention X is performed on a certain user. That is, the diffusion model generation unit 25 generates a network diffusion model M that reflects information indicating the user influence of each user on the false information c obtained based on social media data, the propagation probability of the false information c spreading from one user to another, and the intervention effect on each user in the social network.
[0084] The diffusion model generation unit 25 may construct a network diffusion model M, which is, for example, an extension of the independent cascade model, based on the estimation results from the individual intervention effect estimation unit 241, the user influence estimation unit 242, and the propagation probability estimation unit 243. The network diffusion model M is a mathematical model that describes the diffusion of misinformation c on a social network. For example, if misinformation c propagates between user u and user v with a propagation probability p(u, v, c), user u who receives the misinformation c will share the misinformation c with a probability indicated by user influence s(u, c) and share a rebuttal to the misinformation c with a probability of 1-s(u, c). In this case, if intervention X is performed on user u, the network diffusion model M estimates the propagation probability p of misinformation c from user u to another user when intervention X is performed on user u, based on the result that user influence of user u has decreased from s(u, c) to s(u, c)-ε(u, X).
[0085] The target selection unit 26 selects users who are targets of intervention X based on the network diffusion model M. For example, the target selection unit 26 selects pre-banking targets (intervention target users) who minimize the spread rate of misinformation c.
[0086] Generally, it is not economically realistic to intervene with all users of social media. Therefore, when intervening with users on social media, only a portion of the total users can be intervened. Therefore, under the network diffusion model M, the target selection unit 26 selects k intervention target users by simulation or approximation algorithm, so as to reduce the prevalence rate of misinformation c (the proportion of users deceived by misinformation c) when intervention X is performed. The value of k is determined taking into account factors such as the cost of intervention.
[0087] Specifically, the target selection unit 26 formulates a combinatorial optimization problem that minimizes the expected value of the penetration rate on a given social network, and selects intervention target users according to some approximation algorithm.
[0088] (1) Example of formulating a combinatorial optimization problem to minimize the expected value of the diffusion rate In a social network G = (V, E), the expected value of the diffusion rate of misinformation c when intervention is performed on a user set S ⊂ V under a network diffusion model M is defined as σ(S). Then, the number of elements in the user set S is determined to minimize σ(S). * Let us consider a combinatorial optimization problem to identify the target users S. The Vs that make up the social network G are nodes (circles in Figure 8-1), and Es are edges (lines connecting circles in Figure 8-1). * is expressed by the following formula (Q): * =argmin σ(S) subject to |S|=k...(Q)
[0089] (2) Example of Approximation Algorithm for Solving the Optimization Problem of the Above Formula (Q) As a simple example, the target selection unit 26 solves the problem of finding k users who should intervene before information is transmitted, as shown in Formula (Q), in order to minimize the expected value σ(S) from the user set S. As an approximation algorithm for this, the target selection unit 26 may calculate the expected value σ({u}) of the penetration rate when intervention is performed only for user u, for all users u in the user set S, and select k users as intervention target users in ascending order of σ({u}).
[0090] Or, S current For all users u not included in current The expected value of the penetration rate when intervention is performed only on users included in current ) and calculate the expected value σ({u} ∪ S current ) is the user for which current The operation of sequentially adding to current This is continued until the number of elements of becomes k, and finally S current It is also possible to determine k users included in S as intervention target users. The above procedure is known as the greedy method. current denotes the set of intervening users in the current iteration.
[0091] One possible method for calculating the expected value σ(Z) of the penetration rate when intervention is performed on a certain set Z is to "perform numerical simulations multiple times and calculate the average value." Another possible method for calculating the expected value σ(Z) of the penetration rate is to "approximate the structure of a given social network G with a simple graph structure such as a tree structure and theoretically derive the expected value."
[0092] The result output unit 27 displays the selection results, such as a list of selected intervention target users. The result output unit 27 processes the selection results into a format that supports intervention decisions by countermeasure actors (social media, news media, specialized institutions, etc.) who take measures against the spread of misinformation, and outputs the results to the countermeasure actors. The estimation device 20 processes the estimation results into formats such as a list of users who contribute significantly to the intervention effect, or a depiction of the simulation results of the intervention effect when intervention is performed on specified users. Using the estimation results processed in this way can help countermeasure actors design user intervention strategies for efficient countermeasures against the spread of misinformation.
[0093] <Estimation Process> The processing procedure of the estimation process according to the second embodiment will be described with reference to Fig. 7 . This estimation process is executed by, for example, the estimation device 20. Fig. 7 is a flowchart showing the processing procedure of the estimation process according to the second embodiment. The data collection unit 21 collects social media data from social media or the like (step S11). The data collection unit 21 converts the social media data into a format that can be processed by the estimation unit 24, and then stores the converted data in the social media DB 22. The data collection unit 21 may also collect user survey data 33.
[0094] Next, the individual intervention effect estimation unit 241 analyzes the correlation between the effect amount of intervention X on a certain user and user factors based on the user's social media data (step S12). The individual intervention effect estimation unit 241 may collect social media data and / or user survey data 33 and analyze the correlation between user factors and the intervention.
[0095] Next, the individual intervention effect estimation unit 241 estimates an individual intervention effect for a certain user based on the correlation between the user factor and the intervention X using the user's social media data (step S13). For example, if the effect amount of the intervention X correlates with the user factor A and the correlation coefficient at that time is r, the individual intervention effect estimation unit 241 may calculate the individual intervention effect ε using the above formula (P).
[0096] Next, the user influence estimation unit 242 estimates the user influence s(u, c) of a certain user u on the false information c (step S14).
[0097] Next, the propagation probability estimation unit 243 estimates the propagation probability p(u, v, c) that the false information c shared by a user u will propagate to another user v, based on the user characteristics and past diffusion data of the false information c (step S15).
[0098] Next, the diffusion model generation unit 25 uses the estimation result of the estimator 24 to generate a network diffusion model M that estimates the propagation probability p of the misinformation c from user u to user v when intervention X is performed (step S16). For example, when intervention X is performed on user u before the dissemination of the misinformation c, the diffusion model generation unit 25 generates a network diffusion model M that estimates the propagation probability p of the misinformation c from user u to user v when intervention X is performed, using the estimated results of the individual intervention effect ε(u, X), the user influence s(u, c), and the propagation probability p(u, v, c) that the misinformation c shared by user u propagates to another user v.
[0099] Next, the target selection unit 26 selects k users who should intervene before the information is transmitted based on the network diffusion model M (step S17), where k is an integer equal to or greater than 1. For example, the target selection unit 26 uses equation (Q) to select k users from the user set S who should intervene before the information is transmitted in order to minimize the expected value σ(S).
[0100] The result output unit 27 outputs the selection result, such as a list of selected intervention target users (step S18). The result output unit 27 processes the selection result in a form that will support the intervention decision of the countermeasure person who takes measures against the spread of misinformation, outputs the result to the countermeasure person, and ends the process.
[0101] <Effects of the Second Embodiment> By using the estimation results of the estimation device 20, a countermeasure provider can efficiently select users for whom intervention will be highly effective, knowing what kind of intervention should be performed on whom to effectively suppress the spread of misinformation. Figures 8-1 to 8-3 are diagrams illustrating examples of intervention effects.
[0102] Figures 8-1 to 8-3 show an example of a social network G = (V, E). As shown in Figure 8-1, V represents nodes 511 to 524. E represents the lines connecting the nodes 511 to 524. In Figures 8-1 to 8-3, nodes 511 to 524 indicated by white circles (◯) represent users who do not share information. Node V includes a node 511 for user o, a node 513 for user u, and a node 521 for user v.
[0103] Figure 8-2 shows an example of a diffusion network for misinformation c when user o disseminates misinformation c without implementing prebanking intervention on social media. Nodes 511 to 516, 518, 519, and 522 marked with diagonal lines in Figure 8-2 represent users who shared the misinformation c. Nodes 521, 523, and 524 marked with black circles in Figure 8-2 represent users who shared the true information. Nodes 517 and 520 marked with white circles in Figure 8-2 represent users who did not share the information.
[0104] User v at node 521 who receives false information c will refute the false information c and share the true information with a probability of 1-s(v, c), where s(v, c) is the user influence of user v. As a result, if prebanking intervention is not implemented and user o sends false information c, the false information c will be shared with nine users at nodes 511 to 516, 518, 519, and 522.
[0105] FIG. 8-3 shows an example of a diffusion network for misinformation c when a prebanking intervention is implemented on social media and user o disseminates misinformation c. When implementing a prebanking intervention on social media, a prebanking target is efficiently selected using the estimation method according to the second embodiment. For example, assume that user u is selected as the prebanking target and prebanking is provided to user u at node 513. In this case, by providing prebanking to user u, user u's user influence decreases from s(u, c) to s(u, c)-ε(u, X), which leads user u at node 513 to refute the misinformation c. As a result, the spread of misinformation c from user u at node 513 to users at nodes 514 and 515 is suppressed, and true information is propagated from user u at node 513 to users at nodes 514 and 515. As a result, if pre-banking intervention is performed on user u and user o transmits false information c, the false information c will be shared with six users, nodes 511, 512, 516, 518, 519, and 522. Note that nodes 517 and 520, which are shown with white circles (◯) in Figure 8-3, indicate users with whom information is not shared.
[0106] In the estimation method according to the second embodiment, the estimation device 20 constructs a specific network diffusion model M (mathematical model) that reflects factors such as a user's susceptibility to misinformation, the probability of misinformation spreading, and the effect of individual intervention on each user. The estimation device 20 then selects intervention target users who will maximize the intervention effect under the network diffusion model M. As described above, when implementing a prebanking intervention on social media, the effect of prebanking (the effect of suppressing the spread of misinformation) can be maximized by efficiently selecting prebanking targets using the estimation method according to the second embodiment.
[0107] <Method for building a learning model when there is little social media data> In the above, we estimated each user's misinformation sharing behavior and the intervention effect of anti-misinformation technology on social networks based on data collected from social media.
[0108] However, because the type and / or amount of content distributed and the data that can be collected vary depending on the social media, it may be difficult to secure sufficient data for highly accurate estimation using conventional techniques on some social media platforms. For example, on emerging social media platforms, there are many users who have few opportunities to engage with misinformation content, so it may be possible to collect only a small amount of training data necessary to estimate users' misinformation sharing behavior. In such cases, it may be desirable to use data on the results of intervention X implemented on social media platform A to estimate the intervention effect if intervention X were implemented on social media platform B, where intervention X has not yet been implemented.
[0109] Therefore, we will build an estimation system that transfers the knowledge of a learning model that estimates users' misinformation sharing behavior and intervention effects on social media A to another social media B, which has a lack of data, and uses it for learning.
[0110] An estimation system 100 that executes a learning model construction method when there is little data on social media B will be described with reference to FIGS. 9, 10-1, and 10-2. FIG. 9 is a diagram schematically illustrating an example of the configuration of the estimation system 100 according to an embodiment. FIG. 10-1 is a flowchart illustrating the processing procedure for data preprocessing according to an embodiment. FIG. 10-2 is a flowchart illustrating the processing procedure for data learning processing according to an embodiment.
[0111] 9 , the estimation system 100 includes an estimation device 10A, an estimation device 10B, a multi-platform data pre-processing unit 16, and a transfer model training unit 17. The estimation system 100 has a learning model knowledge transfer function due to the multi-platform data pre-processing unit 16 and the transfer model training unit 17.
[0112] The configurations of the estimation devices 10A and 10B in Fig. 9 are the same as the configurations of the estimation device 10 in Fig. 1, and the data collection units 11A and 11B correspond to the data collection unit 11 of the estimation device 10. The data processing units 12A and 12B correspond to the data processing unit 12 in Fig. 1. The social media A database (DB) 22A and the social media B database (DB) 22B are databases (DBs) that aggregate data of social media A or social media B, the user activity data 141, the user characteristic data 142, and the social network data 143 in Fig. 1.
[0113] The user misinformation sharing behavior estimation units 144A and 144B correspond to the user misinformation sharing behavior estimation unit 144 in Fig. 1. The network intervention effect estimation units 145A and 145B correspond to the network intervention effect estimation unit 145 in Fig. 1. The estimation result output units 15A and 15B correspond to the estimation result output unit 15 in Fig. 1. Countermeasure implementers and end users can use a browser to check the estimation results, such as a list of intervention target users, to help design a user intervention strategy for efficient countermeasures against the spread of misinformation.
[0114] The multi-platform data preprocessing unit 16 extracts features common to the source social media A and the destination social media B, or selects appropriate features, and corrects and converts them to adapt to the transfer by domain adaptation or feature conversion.
[0115] The transfer model training unit 17 uses the features extracted and selected by the multi-platform data preprocessing unit 16 to train a learning model that estimates users' misinformation sharing behavior and intervention effects on social media A.
[0116] A part of the trained model (such as the feature extraction layer and weight parameters) is transplanted into the training model in the user misinformation sharing behavior estimation unit 144B and the network intervention effect estimation unit 145B in social media B, and the training model is retrained and fine-tuned using a small amount of data from social media B.
[0117] Using the above learning model, the user misinformation sharing behavior estimation unit 144B and the NW intervention effect estimation unit 145B of the social media B of the estimation device 10B estimate the user misinformation sharing behavior and the NW intervention effect.
[0118] This makes it possible to transfer knowledge from a learning model that estimates users' misinformation sharing behavior and intervention effects on social media A to another social media B that lacks data and use it for learning. Below, we will specifically explain the processing procedures for data preprocessing and data learning processing when transferring knowledge from a learning model that estimates users' misinformation sharing behavior and intervention effects on social media A to social media B that lacks data and use it for learning.
[0119] <Data Preprocessing> Figure 10-1 shows the processing procedure for data preprocessing according to an embodiment. The processing content of the data preprocessing is to preprocess data from both social media A (hereinafter also referred to as "SM A") and social media B (hereinafter also referred to as "SM B") in order to transfer knowledge from social media A to social media B. The input data for the data preprocessing is the collected data from "SM A" and "SM B", and the output data is the preprocessed data from "SM A" and "SM B".
[0120] In the data preprocessing, the multi-platform data preprocessing unit 16 compares the collected data of "SM A" and "SM B" and extracts common or similar features (step S21). As an example of this processing, the multi-platform data preprocessing unit 16 may create a list of conceptually or formally common features between "SM A" and "SM B" (e.g., "number of reposts" and "number of shares") and extract the features listed in the list. As another processing example, the multi-platform data preprocessing unit 16 may score the association between the features of "SM A" and "SM B" using correlation analysis or mutual information and extract highly similar features (e.g., features above a threshold or the top x% of features).
[0121] Next, the multi-platform data preprocessing unit 16 converts the extracted features using domain adaptation or embedding techniques so that the distributions and characteristics of the features match (step S22). As an example of this processing, the multi-platform data preprocessing unit 16 may measure the distributions (e.g., mean and variance) of common or similar features of "SM A" and "SM B," and if the distributions differ significantly, correct them so that the distributions match. As another example, the multi-platform data preprocessing unit 16 may map the features of "SM A" and "SM B" into a common embedding space using embedding techniques such as a Word2Vec model or BERT.
[0122] <Data Learning Process> Fig. 10-2 shows the processing procedure of the data learning process according to the embodiment. The processing content of the data learning process is to transfer the knowledge of the user misinformation sharing behavior estimation model and intervention effect estimation model in "SM A" to the learning model of "SM B". The input data of the data learning process is the preprocessed data of "SM A" and "SM B" in Fig. 10-1, and the output data is the learned model of "SM B".
[0123] In the data learning process, the transfer model training unit 17 generates a source domain model using the preprocessed data of "SM A" and converts it into a reusable format (step S31). As an example of this process, the transfer model training unit 17 may design a learning model using a logistic regression model or a neural network model, train it using the preprocessed data of "SM A", and save the learned parameters.
[0124] Next, the transfer model training unit 17 generates a target domain model using the source domain model (step S32). As an example of this process, in the case of a neural network model, the transfer model training unit 17 may use the feature extraction layer of the trained model from the previous step as is and design only a new output layer. As another example of process, in the case of a logistic regression model, the transfer model training unit 17 may design a regression equation by adding features (variables) specific to "SM B" to the trained model from the previous step.
[0125] Next, the transfer model training unit 17 retrains the target domain model using the preprocessed data of "S M B" (step S33). As an example of this processing, in the case of a neural network model, the transfer model training unit 17 may fix the parameters of the feature extraction layer using the preprocessed data of "S M B" and retrain only the parameters of the output layer. As another processing example, in the case of a logistic regression model, the transfer model training unit 17 may retrain only the regression coefficients of variables specific to "S M B".
[0126] This allows the estimation device 10B to improve the accuracy of estimating users' misinformation sharing behavior and intervention effects, even on social media B where sufficient data cannot be secured, thereby enabling the selection of more appropriate intervention targets. It also reduces the costs required for data collection, reconstruction of learning models, and data labeling.
[0127] The above describes a learning model construction method executed by the estimation device 10 (estimation devices 10A and 10B) according to the embodiment when there is little social media data. However, this learning model construction method can also be applied to the estimation device 20 according to the second embodiment.
[0128] Although the present invention has been described above as an embodiment, the present invention is not limited to the descriptions and drawings that form part of the disclosure of the present invention. In other words, other embodiments, examples, and operational techniques that can be made by those skilled in the art based on the present invention are all included in the scope of the present invention.
[0129] DESCRIPTION OF SYMBOLS 10 Estimation device 11 Data collection unit 12 Data processing unit 13 Estimation processing unit 14 Estimation unit 15 Estimation result output unit 16 Multi-platform data pre-processing unit 17 Transfer model training unit 31 User data 32 Prediction result 141 User activity data 142 User characteristic data 143 Social network data 144 User misinformation sharing behavior estimation unit 145 NW intervention effect estimation unit 20 Estimation device 21 Data collection unit 22 Social media DB 23 Estimation processing unit 24 Estimation unit 25 Diffusion model generation unit 26 Target selection unit 27 Result output unit 33 User survey data 241 Individual intervention effect estimation unit 242 User influence estimation unit 243 Propagation probability estimation unit
Claims
1. An estimation device comprising: a first estimation unit that estimates each user's misinformation sharing behavior using user data acquired based on social network data; and a second estimation unit that builds a social network of friendships from the user's friend data and estimates the intervention effect in the social network based on each user's misinformation sharing behavior.
2. The estimation device described in claim 1, characterized in that the first estimation unit uses a machine learning model to estimate what kind of misinformation-sharing behavior a target user will exhibit when exposed to misinformation and / or how misinformation-sharing behavior will change when the target user receives intervention, based on user activity data indicating the past activities of each user and / or user characteristic data indicating the interests and political bias of each user.
3. The estimation device described in claim 1, characterized in that the second estimation unit estimates, based on the misinformation sharing behavior of each user, how much misinformation can be reduced in the social network and / or how the prevalence rate of the misinformation will change if intervention is made with a specific user or user group, depending on the specific user or user group that made the intervention.
4. An estimation method executed by an estimation device, comprising: a step of estimating the misinformation sharing behavior of each user using user data acquired based on social network data; and a step of constructing a social network of friendships from the user's friend data and estimating the intervention effect in the social network based on the misinformation sharing behavior of each user.
5. An estimation device comprising: an estimation unit that estimates each user's susceptibility to misinformation obtained based on social media data, the propagation probability of the misinformation from each user to other users, and the effect of intervention on each user in a social network; a generation unit that uses the results estimated by the estimation unit to generate a network diffusion model that estimates the propagation probability of the misinformation when intervention is performed; and a selection unit that selects users who will be targets of the intervention based on the network diffusion model.
6. The estimation device described in claim 5, characterized in that the estimation unit estimates the effect of intervention on each user based on the social media data in accordance with the correlation between each user factor regarding the misinformation and the intervention.
7. An estimation method executed by an estimation device, comprising: a step of estimating the degree of influence of each user on misinformation obtained based on social media data, the propagation probability of the misinformation spreading from the user to other users, and the intervention effect in a social network; a step of generating a network diffusion model using the estimated results to estimate the propagation probability of the misinformation when intervention is carried out; and a step of selecting users to be targets of the intervention based on the network diffusion model.
Citation Information
Patent Citations
Method for determining information propagation model factors based on multi-layer weighted network
CN116228448A
SNS(social network service) score utilization system
JP2016115089A
Generating method, generating program and information processing device
JP2023119197A