Propagation prediction method of false and true guiding topic game based on user psychological income and generative adversarial network

The entropy weight method and Simcse model are used to quantify user psychological benefits, and use adversarial generation network and evolutionary game theory to predict the dissemination situation of false and real topics, solving the problem of game relationship prediction and data sparse data, and achieving effective control over the dissemination of false topics.

CN119939022AActive Publication Date: 2025-05-06CHONGQING UNIV OF POSTS & TELECOMM
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202411924373.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-06
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively predict the game relationship between false guiding topics and real topics, and the sparse data and imbalance issues affect the research on the communication trend.

Method used

The entropy weight method is used to quantify user psychological benefits, and the Simcse text representation learning method is used to mine user historical topic data. At the same time, data augmentation network is used to enhance data, and evolutionary game theory is introduced, combining adversarial game scenarios between false and real topics to predict user behavior and topic dissemination trends.

Benefits of technology

Effectively quantify user psychological benefits, improve the quality and diversity of false and real topic data, accurately predict user behavior and topic dissemination trends, and help governments and relevant departments quickly control the spread of false topics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939022A_ABST
    Figure CN119939022A_ABST
Patent Text Reader

Abstract

The invention relates to a propagation prediction method for a false and true guiding topic game based on user psychological income and a generative adversarial network, and belongs to the technical field of Internet application. The method comprises the following steps: defining related parameters in a network topic propagation process; constructing false and real guiding topic data spaces according to the user basic information, the user relation network, the user historical behaviors and the guiding topic data of the user; quantifying the psychological income of the user through Simcse and an entropy weight method, and enhancing the topic feature data of the user through a generative adversarial network; quantifying an adversarial game relationship between the false topic and the real topic through an evolutionary game theory, and fusing the mutual influence of the false topic and the real topic into a user relationship network adjacency matrix; and in combination with the enhanced user topic feature data and an adjacent matrix containing mutual influence information, a PI-GCN model is adopted to predict user behaviors, and the propagation trend of future guiding topics is predicted and counted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of Internet applications, and relates to a method for predicting the spread of false and real guiding topic games based on user psychological benefits and adversarial generation networks. Background Art

[0002] In recent years, many domestic and foreign scholars have conducted in-depth research on false and misleading topic situation awareness and have achieved certain results. Their research mainly focuses on the following three aspects: factors related to topics and users, sparse and unbalanced data, and the influence relationship between users. This section focuses on the three issues mentioned above and explains the results achieved by scholars in recent years.

[0003] First, in terms of studying the factors associated with topics and users, Jing et al. constructed a user dissemination intention inference model based on dissemination characteristics (behavioral characteristics and time characteristics) and a bidirectional back propagation (B-BP) deep neural network to discover the relationship between false guiding topics and the intensity of users' willingness to spread topics. L. Fang et al. proposed a novel tree variational autoencoder model and used a cross alignment method to align multiple modalities to obtain the relationship between rumors and the text's own emotions. D. Varshney et al. proposed a Bayesian network-based method to predict the dissemination trend of potential topic information driven by factors such as user interests and content similarity. The above studies explore the dissemination factors of such false guiding topics by studying the relationship between the characteristics of the topic itself and the user, but few people have considered the role of the psychological benefits obtained by users in participating in the dissemination of false guiding topics in promoting the dissemination of topics.

[0004] Secondly, in response to the problem of data sparsity and imbalance in the study of topic propagation trends, some scholars have tried to use representation learning methods to alleviate the problems caused by data sparsity. Chen J, Kong et al. alleviated the problem of data sparsity in recommendation systems through network representation learning. Fu et al. proposed a new representation learning framework HIN2Vec for heterogeneous information networks to alleviate the problem of inaccurate prediction caused by data sparsity in link prediction applications. The literature uses an activation maximization method - PPGNs (plug and play generative networks) to improve sample quality and sample diversity. Li Q et al. use the advantage of tensor completion in high precision when recovering lost data to perform homomorphic compensation on topic data to solve the sparsity and imbalance of data encountered in the process of rumor propagation. These studies have effectively alleviated the problems caused by data sparsity, but in the topic space structure of the game between false and real topics, no one has considered the impact of data sparsity and imbalance on topic propagation research.

[0005] Finally, the issue of influence relationship between users is studied. Some scholars start with the influence between users, model and study the structural characteristics in social networks, and measure the influence of different user nodes on information dissemination. M. Cha et al. used data from Twitter to compare the three main parameters that affect node influence: in-degree, number of reposts, and number of @s. The study found that nodes with higher influence play an important role in multiple topics. Yuan Fuyong et al. proposed a user influence index model from the perspectives of users and user microblogs. The literature uses a large-scale field experiment to test the role of social networks in online information dissemination. The experiment shows that more weak relationships are the reason for spreading information. Among them, some scholars use game theory to study the interaction between users in social networks and then predict individual behavior. Wang E K et al. proposed an incentive evolutionary game model for opportunity social networks, which effectively combines the credit-based incentive method with the evolutionary game model to eliminate abnormal nodes and keep network nodes stable and reliable. These studies all use evolutionary game theory to study user relationships in social networks, but under the social network space structure of false topics-real topics, the game competition between false topics and real topics and the potential relationship between the influence of users are worthy of researchers' attention and research.

[0006] False and misleading topics refer to inaccurate and false topics that are artificially created by the creator to deliberately mislead people in order to achieve a certain purpose. The content of this type of topic has a certain superficial or one-sided statement, which cannot objectively reflect the essence of things. It is usually purposeful, misleading and verifiable. The rise of some extremely popular social media (such as Facebook, Twitter, Weibo, etc.) has also provided a suitable environment for the spread of false and misleading topics, which has brought serious negative impacts on the daily life of modern humans and the security and stability of the country and society. In the academic community, S / N has published a number of articles on the spread of false and misleading topics in recent years, describing the many hazards caused by the spread of false and misleading topics in social networks. It can be seen that it is very important to effectively use scientific and technological methods to help the government or relevant departments to timely discover and control the spread of false and misleading topics, create a correct and real network topic environment, and spread accurate and objective real and misleading topic information.

[0007] In many online topic propagation events, we found a common phenomenon. Whenever a false topic is widely spread in social networks, the user nodes that spread the false topic are largely because they have forwarded or published related similar topics before and received strong feedback from other netizens. Here is an example to illustrate the research scenario of this embodiment. A user previously forwarded the topic "Persimmons are prone to cause allergies and other adverse reactions" and received many likes, comments, and forwarding. Later, when he saw the unscientifically proven false topic "Persimmons and seafood will cause poisoning", he was obviously more inclined to recognize this information and chose to forward it. Therefore, the false topic also took advantage of the psychological benefits obtained by the user's participation in the relevant topic, and continued to guide the user to forward this false topic. At the same time, in the process of false topic propagation, some incidental topics will be generated, especially the real topics that are opposite to the content of the false topic. This type of real topic is often an accurate and real topic speech proposed by some official media or authoritative agencies to compete for the relevant false topics, such as the real topic "Persimmons and seafood will cause poisoning without scientific basis" proposed by the corresponding authoritative agency. The two topics constitute a game relationship of confrontation and promotion of "false and real". Therefore, starting from the user's psychological benefits and taking into account the game relationship between false and real topics, we can more comprehensively discover the propagation rules of such false and real guiding topics, thereby helping the government, official platforms and other platforms to quickly and accurately control the spread of false topics. In summary, although there are many readers at home and abroad who have made great achievements in the propagation prediction model of such false guiding topics, there are still some challenges in the study of the propagation trend of multi-information false topics from the perspective of topic guidance:

[0008] 1. The impact of the spread of guiding topics on the psychological benefits of users is difficult to quantify. Whether it is a false topic or a real topic, it often has a certain degree of user psychological guidance. It will use the psychological benefits obtained by users from participating in the content related to the topic to further achieve the purpose of letting users spread the guiding topic. How to effectively find the possibility of users being guided from the user's historical data is a challenge. In addition, the feedback obtained by users in the process of participating in the spread of false guiding topic content is multifaceted, and how to effectively extract and calculate the user's psychological benefits is also a difficulty.

[0009] 2. The actual effective data in the space of false and real guiding topics is sparse and unbalanced. The proportion of forwarding of false and real guiding topics in the entire topic propagation space is actually very small. Moreover, after the forwarded users find that they are spreading false topics, they may delete the data of the forwarded false topics, resulting in sparse and unbalanced data, which brings challenges to the research on the propagation rules of false and real guiding topics.

[0010] 3. Measuring the adversarial game relationship between false topics and real topics is also a challenge. In the process of guiding users to participate in their dissemination, false topics often encounter obstacles from real topics. How to truly quantify this adversarial game relationship and effectively measure the impact of users' psychological benefits from participating in topics are issues that need to be seriously addressed. In addition, this type of irregular social network data structure puts forward certain requirements for the selection of prediction models. Summary of the invention

[0011] In view of this, the purpose of the present invention is to provide a propagation prediction method for false and real guiding topic games based on user psychological benefits and adversarial generative networks. The entropy weight method is used to quantify the psychological benefits of users participating in topics, and the Simcse text representation learning method is combined to mine the impact of user-related historical topic data on user psychological benefits. At the same time, adversarial generative networks are used to homomorphically enhance the sample data of false and real guiding topic spaces, and evolutionary game theory is introduced. Combined with the actual scenes of adversarial games in the process of false and real guiding topic propagation, user behavior in the topic space is predicted and studied.

[0012] In order to achieve the above object, the present invention provides the following technical solutions:

[0013] A method for predicting the spread of false and real guiding topics based on user psychological benefits and adversarial generative networks, the method comprising the following steps:

[0014] S1. Define the relevant parameters in the process of network topic propagation and formalize the problem of propagation prediction;

[0015] S2, constructing a false and real guiding topic data space based on the user's basic information, user relationship network, user historical behavior and guiding topic data;

[0016] S3, quantify the user's psychological benefits through Simcse and entropy weight method, and enhance the user's topic feature data through adversarial generation network;

[0017] S4. Quantify the adversarial game relationship between false and real topics through evolutionary game theory, and integrate the mutual influence of false and real topics into the user relationship network adjacency matrix;

[0018] S5. Combining the enhanced user topic feature data and the adjacency matrix containing mutual influence information, the PI-GCN model is used to predict user behavior and predict the spread trend of future guiding topics.

[0019] Further, in step S1, the following parameters are defined:

[0020] Define the user U who participates in the topic t Topic Communication Network Among them, U t represents the set of users participating in the topic at time t, E U represents the set of directed relationship edges between users participating in the topic, e i,j =(u i ,u j )∈E U Represents user u i Follow user u j ;

[0021] Define the user's basic attributes UserAttr(u i ):

[0022] UserAttr(u i )=[age(u i ),gender(u i ),follow(u i ),fans(u i )]

[0023] Where age(u i ) represents user u i Age, gender(u i ) represents user u i Gender, follow(u i ) represents user u i The number of followers, fans(u i ) represents user u i The number of users who follow you;

[0024] Define user activity UserActivation(u i ):

[0025] UserActivation(u i )=α*PublishCount(u i )+β*ForwardCount(u i )

[0026] Among them, PublishCount(u i ) represents user u i The number of original topics published in a period of time, ForwardCount(u i ) represents user u i The number of topics forwarded by others in a period of time, α and β represent the weights of the two factors;

[0027] Define user psychological benefits PsychBenefits (u i ,T):

[0028]

[0029] Among them, CosineSimlarity(T i ,T) is the user's historical topic T i The similarity between the false and real topics T, ActionHeat(T i ) represents the topic T i The comprehensive popularity in three user behavior dimensions, including likes, comments, and reposts; M represents the number of historical topics of the user;

[0030] Define user intimacy UserIntimacy(u i ,u j ,t):

[0031]

[0032] Among them, A i,j is the indicator function, indicating the potential user u i Whether to follow adjacent user u j ; K represents user u i ,u i The total number of interaction topics between ka Represents the user interaction behavior coefficient, which is determined according to different interaction behaviors; is the time decay function, t represents the current time, t ka Represents user u i For user u j The time of the a-th action published on the k-th topic;

[0033] Define the topic propagation heat TopicHeat(T i ):

[0034] TopicHeat(T i )=[Momentary(T i ),Earlier(T i ),AllNums(T i )]

[0035] Among them, Momentary(T i ) represents the number of forwardings of the current topic within the first preset propagation time after its release. i ) indicates the number of reposts of the current topic within the second preset propagation time after it was published. AllNums(T i) indicates the total number of reposts of the current topic;

[0036] Formalize the propagation prediction problem:

[0037]

[0038] Among them, G = {U, E} is the user propagation network of the current false and real topics, which is a directed graph consisting of user nodes U and user relationship edges E, and P = {(p,u i )|u i ∈U, p∈α} represents the attribute set of all users, where α represents the basic attribute set of the user; A={(a,u i ,t)|u i ∈U,a∈β} represents user u i Perform behavior a at time t, where β represents the historical behavior set of the user in topic propagation; Method() represents the prediction method of user participation in topic forwarding behavior proposed in this paper; D U Represents the predicted participation behavior of potential users in false and real guiding topics.

[0039] Furthermore, in step S2, the false and real guiding topic data space is based on the user basic information, user relationship network, user historical behavior and guiding topic data to calculate the user basic characteristics, user relationship adjacency matrix, user own factors and topic driving factors.

[0040] Further, in step S3, the following steps are specifically included:

[0041] S31. Obtain word embedding representation based on contrastive learning through the Simcse model, obtain the corresponding topic representation vector, and calculate the cosine similarity between the historical topic representation vector and the current false and true topic representation vectors;

[0042] S32, calculating the topic heat index weight by entropy weight method;

[0043] S33. Calculate the user psychological benefit based on the calculated text similarity and topic heat index weights.

[0044] Further, in step S31, firstly, a positive and negative sample group (X, X+, X-) is constructed, where X is the original sample, X+ is the positive sample generated by using X, and X- is the negative sample unrelated to X, and the contrast loss is optimized by reducing the distance between X and X+ and increasing the distance between X and X- in the feature space;

[0045] Among them, the InfoNCE loss function is used as the optimization loss function, which is expressed as:

[0046]

[0047] Among them, q, k + and k - are the representation vectors of the original sample, positive sample, and negative sample after normalization, and τ is the temperature hyperparameter;

[0048] The representation vector of historical topic a obtained by the Simcse model is A(a1,a2,...,a i ,...,a n ), the false and true topics b are B(b1,b2,...,b i ,...,b n ), the cosine similarity calculation formula is as follows:

[0049]

[0050] The range of cosine value is [-1,1], -1 means completely dissimilar, 1 means completely similar, the closer to 1, the more similar, and the closer to -1, the less similar.

[0051] Further, in step S32, an objective weight calculation is performed on the comments, likes, and reposts of each topic using the entropy weight method, which includes:

[0052] S321. Standardize the original data using the minimum-maximum standardization method. The standardized data is expressed as:

[0053]

[0054] In the formula, x ij represents the jth sample of the i-th index, x imin and x imax Respectively represent the minimum and maximum values ​​of the i-th indicator in the sample;

[0055] S322. Calculate the proportion of likes, comments, and reposts of each topic in all samples:

[0056]

[0057] Where m is the number of samples;

[0058] S323. Calculate the entropy value of each indicator:

[0059]

[0060] in, is a constant used to normalize the entropy value;

[0061] S324. Calculate the weight of each indicator by entropy value:

[0062]

[0063] Among them, n is the number of indicators.

[0064] Further, in step S33, the calculated text similarity and topic heat index weights are combined to obtain the user psychological benefits of the user's historical topics on the current false and real topic propagation, respectively, as follows:

[0065]

[0066] Where M is the number of historical topics of the user, N is the number of hot indexes of historical topics, and V false and V true The false guiding topic vector and the real guiding topic vector calculated by Simcse respectively.

[0067] Further, in step S4, it includes:

[0068] S41. By introducing the adversarial generative network, we can mine the potential relationship between the false topics and the real topics in the data, and calculate the topic influence Effect (u i ), and quantify the mutual influence between false and real topics through evolutionary game theory:

[0069]

[0070] Among them, Pro false (u i )-Pro true (u i ) are the profit functions of the two game strategies "forwarding false information" and "forwarding true information", Mut false (u i ) and Mut true (u i ) represent the fake topic and real topic after the game for user u i The influence of communication behavior;

[0071] S42. According to the mutual influence of false and real topics under the influence of user psychological benefits, construct the mutual influence adjacency matrix G of false and real topics at time t t :

[0072]

[0073] in,

[0074] Further, in step S41, the adversarial generative network includes a generative model G and a discriminative model D, wherein the goal of the generative model G is to deceive the discriminative model D, and the goal of the discriminative model D is to distinguish the data generated by the generative model G from the real data. The two constitute a dynamic game process, and ultimately the optimization goal is to achieve Nash equilibrium, wherein,

[0075] The real topic feature data set is represented as datas[x1,x2,...,xn]. Assuming that the topic feature sequence obeys a certain distribution P(x), the method of establishing the generative model G is described as finding the maximum likelihood of the generative model, that is:

[0076]

[0077] The generation and discrimination iterative process of topic feature sequence is described as follows: Let Gz represent the topic sample generation model, z represents the data after random sampling of the original topic feature sequence, and the model G generates the randomly sampled data z as topic feature data;

[0078] D is a topic feature sequence discriminant model. For any input feature sequence x, the discriminant model outputs a real number between 0 and 1, which represents the probability that the feature sequence comes from the real collected sample data.

[0079] P datas and P G Represent the distribution of real topic data and generated topic data respectively, then the objective function of the discriminant model D is:

[0080]

[0081] The optimization function of the entire model is expressed as:

[0082]

[0083] The entire optimization process is represented by iterating D and G until the entire process converges. The process is represented as: datas G =GAN(datas),wait for datas G Infinitely close to datas;

[0084] Topic influence Effect (u i ) is composed of user factors and topic driving factors. User factors include: user basic attributes and user activity, namely:

[0085] UserFactors(u i )=UserAttr(u i )*UserActivation(u i )

[0086] The factors driving the topic are: user intimacy and basic attributes of the topic, namely:

[0087] TopicDrive(u i ,T i )=UserIntimacy(u i ,u j ,t)*TopicHeat(T i )

[0088] Taking into account user factors and topic driving factors, the influence function of false topics and the influence function of real topics are constructed using the multivariate linear regression algorithm:

[0089] Effect false (u i )=β0+β1*UserFactors(u i )+β2*TopicDrive(u i ,T false )

[0090] Effect true (u i )=β0+β1*UserFactors(u i )+β2*TopicDrive(u i ,T true )

[0091] Among them, β0, β1, and β2 are partial regression coefficients obtained by training using the multivariate linear regression algorithm, and β1 and β2 represent the proportion of each factor in the influence of the topic;

[0092] According to game theory, two game strategies are defined: "forwarding false information" and "forwarding true information". The profit functions of the two strategies are:

[0093] Pro false (u i )=P1×Effect false (u i )

[0094] Pro true (u i )=P1×Effect true (u i )

[0095] Where P1 and P2 are user u i The proportion of fake topics and real topics spread among friends; Effect false and Effecttrue They are the topic influences under the influence of psychological benefits of related false topics and real topics in the historical topics participated by users;

[0096] Finally, the calculation formula of evolutionary game theory is used, and the impact of user psychological benefits is taken into account to construct the final mutual influence between false and real topics.

[0097] Further, in step S5, PI-GCN is used as the prediction model, wherein the input of the prediction model is:

[0098] Feature matrix X = N × A, where N represents the number of user nodes in the false topic propagation network, and A is the topic feature vector of each user node, which includes the user's own attribute characteristics and topic driving factor characteristics;

[0099] Adjacency Matrix in Game Represents the relationship connection information between all users in the two topic spaces of false topics and real topics under the influence of user psychological benefits at time t;

[0100] First, a two-layer graph convolutional neural network with an intermediate Dropout layer is used as a network rumor forwarding prediction model. First, the weights and biases are randomly initialized. Then, X is multiplied by W and the bias is added, and then multiplied by A^. After that, the RELU function is used as the activation function of this layer, and the Dropout operation is performed during model training. Finally, the SoftMax activation function is used to represent the convolution output as the probability value of different node classifications; it is expressed as:

[0101] Z=f(X,A)=softmax(A^ReLU(A^XW 0 )W 1 )

[0102] Among them, W i is the weight matrix corresponding to the i-th layer network in the graph convolutional network,

[0103] Let the model output Z = P(r,a,d|u i ), the specific definition is as follows:

[0104]

[0105] Among them, if the corresponding D = 1, then the potential user u is judged i The false topic will be forwarded in the next time period; if D = -1, the potential user u i will forward the real topic in the next time period; otherwise, potential user u i In the next period of time, we will not participate in the dissemination of false or true topics.

[0106] The beneficial effects of the present invention are:

[0107] Aiming at the problem of complex influencing factors of users' psychological benefits from participating in topics, the present invention proposes a method for quantifying users' psychological benefits based on Simcse and entropy weight method. The Simcse representation learning method is used to extract and mine historical data related to users' participation in false and real guiding topics, and the relationship between false and real guiding topics and users is found from the aspect of topic content similarity. The entropy weight method is used to perform a multi-indicator comprehensive evaluation on users' feedback on participating in topics, and the psychological benefits of users' participation in false and real guiding topics are calculated.

[0108] Aiming at the problem that effective samples of false and real guiding topics are sparse, the present invention proposes a guiding topic data enhancement method based on a generative adversarial network. The topic samples are homomorphically enhanced through an unsupervised generative adversarial network model, so as to more realistically restore the relationship between the user's psychological benefits and the propagation trends of false and real guiding topics.

[0109] The present invention proposes a false and real guiding topic game propagation model based on PI-GCN, which effectively processes the special structure data of social network, introduces evolutionary game theory, and combines the psychological benefits of users participating in topics, truly quantifies the driving influence of users under the influence of two topic games, constructs the user psychological benefit fluctuation matrix, and finally effectively predicts the user behavior in the entire network topic space through the propagation prediction model.

[0110] The present invention can not only effectively predict user behavior and popularity in the false-true guiding topic space, but also more realistically reflect the propagation trend of such adversarial game guiding topics under the user's psychological benefit factors and user social relationships.

[0111] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0112] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:

[0113] Figure 1 It is a schematic diagram of the overall process of the propagation prediction method of the false and real guiding topic game based on user psychological benefits and adversarial generative network of the present invention;

[0114] Figure 2A schematic diagram of the process of obtaining a topic vector using the Simcse method of the present invention. DETAILED DESCRIPTION

[0115] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0116] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0117] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0118] See also Figure 1-2 , which is a propagation prediction method for false and real guiding topic games based on user psychological benefits and adversarial generative networks.

[0119] Example

[0120] This embodiment provides a method for predicting the spread of false and real guiding topics based on user psychological benefits and adversarial generative networks. Figure 1 As shown, it includes the following steps:

[0121] S1. Define the relevant parameters in the process of network topic propagation and formalize the problem of propagation prediction;

[0122] S2, constructing a false and real guiding topic data space based on the user's basic information, user relationship network, user historical behavior and guiding topic data;

[0123] S3, quantify the user's psychological benefits through Simcse and entropy weight method, and enhance the user's topic feature data through adversarial generation network;

[0124] S4. Quantify the adversarial game relationship between false and real topics through evolutionary game theory, and integrate the mutual influence of false and real topics into the user relationship network adjacency matrix;

[0125] S5. Combining the enhanced user topic feature data and the adjacency matrix containing mutual influence information, the PI-GCN model is used to predict user behavior and predict the spread trend of future guiding topics.

[0126] In step S1 of this embodiment, the main purpose of this method is to predict the development trend of potential users' forwarding of false topics or real topics and false and real guiding topics in the topic space by analyzing the basic information of users, the relationship between users, the psychological benefits of users, and the game relationship between false and real topics. Based on this, this embodiment defines the users participating in the topic, the topic propagation network, the basic attributes of the users themselves, the user activity, the psychological benefits of users, the intimacy of users, and the popularity of topic propagation, and formalizes the problem solved by the present invention according to the defined content.

[0127] Specifically, define the user U participating in the topic t Topic Communication Network Among them U t represents the set of users participating in the topic at time t, E U represents the set of relationship edges between users participating in the topic, and the relationship edge is a directed edge, where e i,j =(u i ,u j )∈E U Represents user u i Followed user u j , that is, user u i Is user u j The set of users participating in the topic and the set of user relationships constitute the topic propagation network

[0128] Define the user's basic attributes UserAttr(u i ), the basic attributes of users include, for example, user age, user gender, number of users followed by users, and number of users followed by users. These attributes have a certain impact on whether users will participate in topic dissemination. Therefore, the basic attributes of users are defined in the text as:

[0129] UserAttr(u i )=[age(u i ),gender(u i ),follow(u i ),fans(u i )]

[0130] Where age(u i ) represents user u i Age, gender(u i ) represents user u i Gender, follow(u i ) represents user u i The number of followers, fans(u i ) represents user u i The number of users who follow .

[0131] Define user activity UserActivation(u i ), the user's activity indicates the user's enthusiasm for participating in the spread of topics. Generally speaking, the more active the user is, the greater the probability of forwarding rumors or refuting rumors. In this embodiment, the user's activity is expressed as:

[0132] UserActivation(u i )=β*PublishCount(u i )+β*ForwardCount(u i )

[0133] Among them, PublishCount(u i ) represents user u i The number of original topics published in the previous month, ForwardCount(u i ) represents user u i The number of topics forwarded by others in the previous month. α and β represent the weights of the two factors, ranging from 0 to 1.

[0134] Define user psychological benefits PsychBenefits (u i ,T), the user psychological benefit is expressed as the similarity between the user's historical topic and the current false guiding topic or the real guiding topic, as well as the topic popularity of the historical topic (number of likes, comments, and reposts), which is defined as follows:

[0135]

[0136] Among them, CosineSimlarity(T i ,T) is the user's historical topic Ti The similarity between the false and true topics T, ActionHeat(T i ) represents the topic T i The comprehensive popularity of the three user behavior dimensions of likes, comments, and reposts. M represents the number of historical topics of the user.

[0137] Define user intimacy UserIntimacy(u i ,u i ,t), user intimacy represents the mutual forwarding, liking and commenting behaviors between users and their friends. Users with more frequent interactions are more likely to forward each other's topics. At the same time, since user intimacy has a strong timeliness, this embodiment introduces a time decay function e t To quantify the impact of user interaction, user intimacy is defined as:

[0138]

[0139] Among them, A i,j is the indicator function, that is: A ij =1 indicates potential user u i Followed neighboring user u j , A ij =0 indicates potential user u i No longer following neighboring user u j ; K represents user u i ,u i The total number of interaction topics between ka represents the user interaction behavior coefficient. Specifically,

[0140]

[0141] A ka =1 indicates user u i Retweeted u j Published topic, A ka =0.7 means user u i Commented u j Published topic, A ka =0.5 means user u i Liked u j The topic of the post.

[0142] is the time decay function, t represents the current time, t ka Represents user u i For user u j The time of the ath action published on the kth topic.

[0143] Define the topic propagation heat TopicHeat(T i ), the number of forwardings of topic information in different time periods represents the spread heat of the topic. Therefore, this embodiment defines the spread heat of the topic as:

[0144] TopicHeat(T i )=[Momentary(T i ),Earlier(T i ),AllNums(T i )]

[0145] Among them, Momentary(T i ) represents the number of reposts of the current topic after 1 hour of propagation since its release. i ) indicates the number of reposts of the current topic after it was published one day ago, while AllNums(T i ) indicates the total number of reposts of the current topic.

[0146] Then, according to the parameters defined above, the problem of the present invention and the input and output of the problem are clarified, wherein the user propagation network of the current false and real topics is defined as a directed graph consisting of user nodes U and user relationship edges E, that is, G = {U, E}; at the same time, this embodiment uses P = {(p,u i )|u i ∈U, p∈α} represents the attribute set of all users, where α represents the basic attribute set of the user; A={(a,u i ,t)|u i ∈U,a∈β} represents user u i Perform behavior a at time t, where β represents the historical behavior set of the user in topic propagation; then, according to the user's historical participation in topics, the psychological benefits of the user's participation in the current topic are mined, and finally, the potential user's participation behavior D in false and real guiding topics is predicted by combining the user's own attributes, user intimacy, and user relationship network U , a clearer problem definition is expressed as:

[0147]

[0148] In step S2 of this embodiment, user basic information, user relationship network, user historical behavior and guiding topic data are used as model inputs to construct a false and real guiding topic data space, and important information such as user basic characteristics, user relationship adjacency matrix, user own factors and topic driving force factors are calculated.

[0149] In step S3 of the present embodiment, the Simcse method is used as the core text representation method for calculating the similarity between the user's historical topics and the false and real topics. In social networks, the spread of topics is often based on some historical behaviors of users. For example, if a user has published or forwarded topic information similar to the false and real topics, then the user is more likely to participate in the spread of the current topic. Therefore, how to effectively calculate the similarity between the user's historical topics and the false and real topics is particularly important. In particular, the present embodiment adopts the Simcse word embedding representation method based on contrastive learning, which can more effectively process short text type data such as user topics. At the same time, Simcse can effectively combine and utilize the Bert pre-training model, and effectively generate accurate topic text vectors in the unsupervised learning mode, which can achieve a more accurate short text similarity calculation effect. Therefore, the present embodiment selects the Simcse method as the core text representation method for calculating the similarity between the user's historical topics and the false and real topics.

[0150] Simcse stands for Simple Contrastive Learning of Sentence Embeddings, which can generate high-quality word vectors through contrastive learning in an unsupervised environment. The idea of ​​contrastive learning is to construct a positive and negative sample group (X, X+, X-), where X is the original sample, X+ is the positive sample generated using X, and X- is the negative sample unrelated to X. The contrast loss is optimized by reducing the distance between X and X+ and increasing the distance between X and X- in the feature space, thereby improving the quality of the obtained feature vector. The positive and negative sample group (X, X+, X-) is the key in contrastive learning. In Simcse, a sentence is input into the neural network twice. Due to the presence of a dropout layer in the model, the neurons will be randomly inactivated, so the feature vectors generated by the two inputs are also different. The vectors obtained by the two inputs are X and X+, while the other sentences in the same batch are X-.

[0151] The Simcse model uses the InfoNCE loss function, whose main goal is to optimize sentence embedding through a contrastive learning mechanism so that semantically similar sentences are closer in the embedding space, while dissimilar sentences are farther apart. The formula of the InfoNCE loss function is as follows:

[0152]

[0153] Among them, q, k + and k -are the representation vectors of the original sample, positive sample, and negative sample after normalization, and τ is the temperature hyperparameter. Since the text representation has been normalized, the form of the InfoNCE loss function is essentially the cross entropy Softmax loss commonly seen in multi-classification tasks. The numerator in the loss function is the similarity between positive sample pairs, and the denominator is the similarity of all samples in the same batch. The training goal of the model is to minimize the loss function, that is, the larger the numerator, the smaller the denominator, the closer the positive sample pairs are in the same data set, and the farther the negative sample pairs are from each other. This is the core idea of ​​contrastive learning. τ is a common hyperparameter in Softmax. The smaller τ is, the closer Softmax is to the real max function, and the larger τ is, the closer it is to a uniform distribution. The overall schematic diagram of the Simcse model is as follows Figure 2 As shown in the figure, for the historical topic word vectors and the false and real topic word vectors calculated by the Simcse model, the cosine similarity method is used to calculate the similarity between false topics, real topics and historical topics. Cosine similarity can be used to judge the similarity between texts. It refers to the cosine value of the angle between two N-dimensional vectors in N-dimensional space. The cosine value is used to measure the difference between two vectors. Assume that the representation vector of historical topic a obtained by the Simcse model is A(a1, a2, ..., a i ,...,a n ), the false and true topics b are B(b1,b2,...,b i ,...,b n ), the cosine similarity calculation formula is as follows:

[0154]

[0155] The range of cosine value is [-1,1], -1 means completely dissimilar, 1 means completely similar, the closer to 1, the more similar, and the closer to -1, the less similar.

[0156] The popularity of a topic is often the core factor that reflects the psychological benefit feedback that users obtain from participating in the topic, and the popularity of the topic that users have forwarded and published in history is reflected in many aspects, including the number of forwarded topics, the number of comments, the number of likes, etc. The entropy weight method is a weight allocation method for comprehensive evaluation of multiple indicators, which is widely used in decision analysis, evaluation and ranking. It does not rely on the subjective judgment of decision makers, and automatically calculates the weight of each indicator through statistical analysis of data. It is very suitable for multi-indicator and multi-dimensional decision-making problems, can comprehensively consider the weight of each indicator, and can automatically adjust the strategy according to the change of data, which is suitable for handling complex decision-making environments. Therefore, this embodiment starts with the topic heat indicators such as the number of forwarded topics, the number of comments, and the number of likes, and adopts the entropy weight method as a comprehensive evaluation method, so as to more objectively and effectively measure the user's historical related topic heat, and then calculate the psychological benefit data of the user's possible participation in false and real topics.

[0157] This embodiment uses the entropy weight method to calculate the objective weights of the three dimensions of comments, likes, and reposts for each topic, so as to calculate the direct psychological benefits of each topic to the user. The specific calculation steps are as follows:

[0158] 1. Data standardization: Standardize the original data, eliminate the influence of dimension, and use the minimum-maximum standardization method. The standardized data can be expressed as:

[0159]

[0160] Among them, x ij represents the jth sample of the i-th index, x imin and x imax They represent the minimum and maximum values ​​of the i-th indicator in the sample respectively.

[0161] 2. Calculate the proportion of each indicator: Calculate the proportion of likes, comments, reposts and other indicators of each topic in all samples:

[0162]

[0163] Among them, m is the number of samples.

[0164] 3. Calculate the entropy value: Calculate the entropy value of each indicator. The calculation formula of the entropy value is:

[0165]

[0166] in, is a constant used to normalize the entropy value.

[0167] 4. Calculate weight: Calculate the weight of each indicator through entropy value. The weight calculation formula is:

[0168]

[0169] Among them, n is the number of indicators.

[0170] Combining the text similarity obtained by the Simcse model and the topic heat index weight calculated by the entropy weight method, we can obtain the user psychological benefits of the user's historical topic on the current false and real topic propagation as follows:

[0171]

[0172] Where M is the number of historical topics of the user, N is the number of hot indicators of historical topics (number of likes, number of reposts, number of comments), V false and V true The false guiding topic vector and the real guiding topic vector calculated by Simcse respectively.

[0173] In step S4 of this embodiment, a generative adversarial network (GAN) is introduced to solve the problem of sparse data of false and real topics, and its adversarial mechanism is used to deeply explore the potential relationship generated by the mutual adversarial generation of false topics and real topics in the data, and generate high-quality sample data based on this relationship.

[0174] GAN is composed of a generative model G and a discriminative model D. The goal of the false and real topic data generation model in this embodiment is to generate real topic data as much as possible to deceive the discriminative model D, and the goal of the discriminative model D is to distinguish the data generated by the generative model D from the collected real data as much as possible, so that G and D constitute a dynamic "game process". The goal of optimization is to achieve Nash equilibrium, that is, the generative model and the discriminative model improve their respective generation and discrimination capabilities in the process of continuous optimization and learning, so that the model can generate data that is homomorphic and distributed with the collected topic samples, thereby generating good user behavior and topic sample data to alleviate the sparsity of the actual effective cross-data in the false and real topic space.

[0175] The topic-related data set used in this embodiment can be expressed as datas[x1, x2, ..., xn]. Assuming that the topic feature sequence obeys a certain distribution P(x), the method for establishing the generative model G can be described as finding the maximum likelihood of this generative model, that is:

[0176]

[0177] The iterative process of generating and discriminating topic feature sequences is described as follows: Let Gz represent the topic sample generation model, z represents the data after random sampling of the original topic feature sequence, and the model G generates the randomly sampled data z as topic feature data. D is a topic feature sequence discrimination model. For any input feature sequence x, Dx will output a real number between 0 and 1 to represent the probability that the feature sequence comes from the real collected sample data. datas and P G Represent the distribution of real topic data and generated topic data respectively, then the objective function of the discriminant model is:

[0178]

[0179] The optimization function of the entire model can be expressed as:

[0180]

[0181] The whole optimization process is represented by iterating D and G until the whole process converges. This process is represented by: datas G =GAN(datas),wait for datas G Infinitely close to datas.

[0182] Topic influence Effect (u i ) is composed of user factors and topic driving factors. User factors include: user basic attributes and user activity, namely:

[0183] UserFactors(u i )=UserAttr(u i )*UserActivation(u i )

[0184] The factors driving the topic are: user intimacy and basic attributes of the topic, namely:

[0185] TopicDrive(u i ,T i )=UserIntimacy(u i ,u j ,t)*TopicHeat(T i )

[0186] Taking into account user factors and topic driving factors, the influence function of false topics and the influence function of real topics are constructed using the multivariate linear regression algorithm:

[0187] Effect false (u i)=β0+β1*UserFactors(u i )+β2*TopicDrive(u i ,T false )

[0188] Effect true (u i )=β0+β1*UserFactors(u i )+β2*TopicDrive(u i ,T true )

[0189] Among them, β0, β1, and β2 are partial regression coefficients obtained by training using the multivariate linear regression algorithm, while β1 and β2 represent the proportion of each factor in the influence of the topic.

[0190] Due to the particularity and complexity of social networks, when a piece of false information is spreading, there is also real information that is competing with it. This game relationship is also an important factor affecting the user's forwarding behavior. Therefore, this embodiment uses evolutionary game theory to quantify the mutual influence of false and real topics. First, two game strategies are defined based on game theory: "forwarding false information" and "forwarding real information". The profit functions of the two strategies are:

[0191] Pro false (u i )=P1×Effect false (u i )

[0192] Pro true (u i )=P1×Effect true (u i )

[0193] Where P1 and P2 are user u i The proportion of fake topics and real topics spread among friends. false and Effect true They are the topic influences under the influence of the psychological benefits of related false topics and real topics in the historical topics that users participated in. Finally, the calculation formula of evolutionary game theory is used, and the influence of user psychological benefits is taken into account to construct the final mutual influence between false and real topics as follows:

[0194]

[0195] Among them, Mut false (u i ) and Mut true (ui ) represent the fake topic and real topic after the game for user u i The influence of communication behavior.

[0196] Finally, according to the mutual influence of false and real topics under the influence of user psychological benefits, the adjacency matrix G of the mutual influence of false and real topics at time t is constructed. t :

[0197]

[0198] in,

[0199] In step S5 of this embodiment, PI-GCN is used as the prediction model of the present invention, and the GCN graph convolutional neural network is used to process typical non-Euclidean structure data such as social networks. Taking into account the promotion (Promotion) and inhibition (Inhibition) relationship between false topics and real topics in the propagation, a group behavior prediction model for false and real information propagation based on PI-GCN is proposed. The goal of the prediction task of this embodiment is to predict the participation of potential user nodes in the false topic. If they participate, it is determined whether they forward a false topic or a real topic, and then it is converted into a three-classification task. The model input of this embodiment is as follows:

[0200] 1. Feature matrix X = N × A, where N represents the number of user nodes in the false topic propagation network, and A is the topic feature vector of each user node, which includes the user's own attribute characteristics and topic driving factor characteristics.

[0201] 2. Adjacency Matrix under Game It represents the relationship connection information between all users in the two topic spaces of false topics and real topics under the influence of user psychological benefits at time t.

[0202] In the application of this embodiment, this embodiment uses a two-layer graph convolutional neural network with an intermediate Dropout layer as a network rumor forwarding prediction model. First, the weights and biases are randomly initialized, then X is multiplied by W and the bias is added, and multiplied by A^. After that, the RELU function is used as the activation function of this layer, and the Dropout operation is performed during model training. Finally, this embodiment uses the SoftMax activation function to represent the convolution output as the probability value of different node classifications. The specific formula is expressed as:

[0203] Z=f(X,A)=softmax(A^ReLU(A^XW 0 )W 1 )

[0204] Among them, W iis the weight matrix corresponding to the i-th layer network in the graph convolutional network,

[0205] Since this system model solves a three-class prediction problem, let the model output Z = P(r, a, d|u i ), the specific definition is as follows:

[0206]

[0207] Among them, if the corresponding D = 1, then the potential user u is judged i The false topic will be forwarded in the next time period; if D = -1, the potential user u i will forward the real topic in the next time period; otherwise, potential user u i In the next period of time, we will not participate in the dissemination of false or true topics.

[0208] In this embodiment, the input of the model includes topic information data, the user relationship network G = {U, E} in the false and real topic space at time t, and the user's own attribute set P = {(p,u i )|u i ∈U,p∈α}, user historical behavior set A={(a,u i ,t)|u i ∈U,α∈β}. The model constructs a false and real mutual influence matrix through evolutionary game theory and user psychological benefits, and uses the GCN graph convolutional neural network to predict the forwarding of false and real guiding topics by potential users at time t+1. The specific propagation model algorithm is shown in Table 1:

[0209] Table 1

[0210]

[0211]

[0212] In addition, this embodiment analyzes the time complexity of this algorithm. The overall time complexity of the user psychological benefit calculation method composed of Simcse and entropy weight method is M×(O(N)+O(1))~O(MN), where M is the number of topics forwarded by each user in history. The time complexity of the user feature data enhancement model based on the adversarial generative network is O(KN), where K is the network depth, which is generally much smaller than the number of samples N. The calculation of user factors and topic driving factors in the algorithm is also O(N), and the time complexity of the PI-GCN prediction model is O(N 2 ), the time complexity of the entire algorithm model is O(MN)+O(KN)+O(N 2 )~O(N 2). Since the actual valid data in the entire false and real topics only accounts for a small part, the input data of the entire PI-GCN model is not large. At the same time, the entire model uses two layers of graph convolution, which ensures that the amount of data involved in model training is within a predictable range. Therefore, the time complexity of the entire algorithm model is O(N 2 ), and it is acceptable.

[0213] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.

Claims

1. A method for predicting the spread of false and real guiding topics based on user psychological benefits and adversarial generative networks, characterized by: The method comprises the following steps: S1. Define the relevant parameters in the process of network topic propagation and formalize the problem of propagation prediction; S2, constructing a false and real guiding topic data space based on the user's basic information, user relationship network, user historical behavior and guiding topic data; S3, quantify the user's psychological benefits through Simcse and entropy weight method, and enhance the user's topic feature data through adversarial generation network; S4. Quantify the adversarial game relationship between false and real topics through evolutionary game theory, and integrate the mutual influence of false and real topics into the user relationship network adjacency matrix; S5. Combining the enhanced user topic feature data and the adjacency matrix containing mutual influence information, the PI-GCN model is used to predict user behavior and predict the spread trend of future guiding topics.

2. According to claim 1, a method for predicting the spread of false and real guiding topics based on user psychological benefits and adversarial generative networks is characterized by: In step S1, the following parameters are defined: Define the user U who participates in the topic t Topic Communication Network Among them, U t represents the set of users participating in the topic at time t, E U represents the set of directed relationship edges between users participating in the topic, e i,j =(u i ,u j )∈E U Represents user u i Follow user u j ; Define the user's basic attributes UserAttr(u i ): UserAttr(u i )=[age(u i ),gender(u i ),follow(u i ),fans(u i )] Where age(u i ) represents user u i Age, gender(u i ) represents user u i Gender, follow(u i ) represents user u i The number of followers, fans(u i ) represents user u i The number of users who follow you; Define user activity UserActivation(u i ): UserActivation(u i )=α*PublishCount(u i )+β*ForwardCount(u i ) Among them, PublishCount(u i ) represents user u i The number of original topics published in a period of time, ForwardCount(u i ) represents user u i The number of topics forwarded by others in a period of time, α and β represent the weights of the two factors; Define user psychological benefits PsychBenefits (u i ,T): Among them, CosineSimlarity(T i ,T) is the user's historical topic T i The similarity between the false and real topics T, ActionHeat(T i ) represents the topic T i The comprehensive popularity in three user behavior dimensions, including likes, comments, and reposts; M represents the number of historical topics of the user; Define user intimacy UserIntimacy(u i ,u j ,t): Among them, A i,j is the indicator function, indicating the potential user u i Whether to follow adjacent user u j ; K represents user u i ,u i The total number of interaction topics between ka Represents the user interaction behavior coefficient, which is determined according to different interaction behaviors; is the time decay function, t represents the current time, t ka Represents user u i For user u j The time of the a-th action published on the k-th topic; Define the topic propagation heat TopicHeat(T i ): TopicHeat(T i )=[Momentary(T i ),Earlier(T i ),AllNums(T i )] Among them, Momentary(T i ) represents the number of forwardings of the current topic within the first preset propagation time after its release. i ) indicates the number of reposts of the current topic within the second preset propagation time after it was published. AllNums(T i ) indicates the total number of reposts of the current topic; Formalize the propagation prediction problem: Among them, G = {U, E} is the user propagation network of the current false and real topics, which is a directed graph consisting of user nodes U and user relationship edges E, and P = {(p,u i )|u i ∈U,p∈α} represents the attribute set of all users, where α represents the basic attribute set of the user; A={(a,u i ,t)|u i ∈U,a∈β} represents user u i Perform behavior a at time t, where β represents the historical behavior set of users in topic propagation; Method represents the prediction method of user participation in forwarding topic behavior; D U Represents the predicted participation behavior of potential users in false and real guiding topics.

3. According to claim 2, a method for predicting the spread of false and real guiding topic games based on user psychological benefits and adversarial generative networks is characterized by: In step S2, the false and real guiding topic data space is based on user basic information, user relationship network, user historical behavior and guiding topic data to calculate user basic characteristics, user relationship adjacency matrix, user own factors and topic driving factors.

4. According to claim 3, a method for predicting the spread of false and real guiding topics based on user psychological benefits and adversarial generative networks is characterized by: In step S3, the following steps are specifically included: S31. Obtain word embedding representation based on contrastive learning through the Simcse model, obtain the corresponding topic representation vector, and calculate the cosine similarity between the historical topic representation vector and the current false and true topic representation vectors; S32, calculating the topic heat index weight by entropy weight method; S33. Calculate the user psychological benefit based on the calculated text similarity and topic heat index weights.

5. The method for predicting the spread of false and real guiding topics based on user psychological benefits and adversarial generative networks according to claim 4 is characterized by: In step S31, firstly, a positive and negative sample group (X, X+, X-) is constructed, where X is the original sample, X+ is the positive sample generated by using X, and X- is the negative sample unrelated to X. The contrast loss is optimized by reducing the distance between X and X+ and increasing the distance between X and X- in the feature space. Among them, the InfoNCE loss function is used as the optimization loss function, which is expressed as: Among them, q, k + and k - are the representation vectors of the original sample, positive sample, and negative sample after normalization, and τ is the temperature hyperparameter; The representation vector of historical topic a obtained by the Simcse model is A(a1, a2, ..., a i , ..., a n ), the false and true topics b are B(b1, b2, ..., b i , ..., b n ), the cosine similarity calculation formula is as follows: The range of cosine value is [-1, 1], -1 means completely dissimilar, 1 means completely similar, the closer to 1, the more similar, and the closer to -1, the less similar.

6. The method for predicting the spread of false and real guiding topics based on user psychological benefits and adversarial generative networks according to claim 5 is characterized by: In step S32, an objective weight calculation is performed on the comments, likes, and reposts of each topic using the entropy weight method, which includes: S321. Standardize the original data using the minimum-maximum standardization method. The standardized data is expressed as: In the formula, x ij represents the jth sample of the i-th index, x imin and x imax Respectively represent the minimum and maximum values ​​of the i-th indicator in the sample; S322. Calculate the proportion of likes, comments, and reposts of each topic in all samples: Where m is the number of samples; S323. Calculate the entropy value of each indicator: in, is a constant used to normalize the entropy value; S324. Calculate the weight of each indicator by entropy value: Where n is the number of indicators.

7. The method for predicting the spread of false and real guiding topics based on user psychological benefits and adversarial generative networks according to claim 6 is characterized by: In step S33, the calculated text similarity and topic heat index weights are combined to obtain the user psychological benefits of the user's historical topics on the current false and real topic propagation, respectively, as follows: Where M is the number of historical topics of the user, N is the number of hot indexes of historical topics, and V false and V true The false guiding topic vector and the real guiding topic vector calculated by Simcse respectively.

8. The method for predicting the spread of false and real guiding topics based on user psychological benefits and adversarial generative networks according to claim 7 is characterized by: In step S4, it includes: S41. By introducing the adversarial generative network, we can mine the potential relationship between the false topics and the real topics in the data, and calculate the topic influence Effect (u i ), and quantify the mutual influence between false and real topics through evolutionary game theory: Among them, Pro false (u i )-Pro true (u i ) are the profit functions of the two game strategies "forwarding false information" and "forwarding true information", Mut false (u i ) and Mut true (u i ) represent the fake topic and real topic after the game for user u i The influence of communication behavior; S42. According to the mutual influence of false and real topics under the influence of user psychological benefits, construct the mutual influence adjacency matrix G of false and real topics at time t t : in, 9. The method for predicting the spread of false and real guiding topics based on user psychological benefits and adversarial generative networks according to claim 8 is characterized by: In step S41, the adversarial generative network includes a generative model G and a discriminative model D, wherein the goal of the generative model G is to deceive the discriminative model D, and the goal of the discriminative model D is to distinguish the data generated by the generative model G from the real data. The two constitute a dynamic game process, and ultimately the optimization goal is to achieve Nash equilibrium, wherein, The real topic feature data set is represented as datas[x1, x2, ..., xn]. Assuming that the topic feature sequence obeys a certain distribution P(x), the method of establishing the generative model G is described as finding the maximum likelihood of the generative model, that is: The generation and discrimination iterative process of topic feature sequence is described as follows: Let Gz represent the topic sample generation model, z represents the data after random sampling of the original topic feature sequence, and the model G generates the randomly sampled data z as topic feature data; D is a topic feature sequence discriminant model. For any input feature sequence x, the discriminant model outputs a real number between 0 and 1, which represents the probability that the feature sequence comes from the real collected sample data. P datas and P G Represent the distribution of real topic data and generated topic data respectively, then the objective function of the discriminant model D is: The optimization function of the entire model is expressed as: The entire optimization process is represented by iterating D and G until the entire process converges. The process is represented as: datas G =GAN(datas), waiting for datas G Infinitely close to datas; Topic influence Effect (u i ) is composed of user factors and topic driving factors. User factors include: user basic attributes and user activity, namely: UserFactors(u i )=UserAttr(u i )*UserActivation(u i ) The factors driving the topic are: user intimacy and basic attributes of the topic, namely: TopicDrive(u i ,T i )=UserIntimacy(u i ,u j ,t)*TopicHeat(T i ) Taking into account user factors and topic driving factors, the influence function of false topics and the influence function of real topics are constructed using the multivariate linear regression algorithm: Effect false (u i )=β0+β1*UserFactors(u i )+β2*TopicDrive(u i ,T false ) Effect true (u i )=β0+β1*UserFactors(u i )+β2*TopicDrive(u i ,T true ) Among them, β0, β1, and β2 are partial regression coefficients obtained by training using the multivariate linear regression algorithm, and β1 and β2 represent the proportion of each factor in the influence of the topic; According to game theory, two game strategies are defined: "forwarding false information" and "forwarding true information". The profit functions of the two strategies are: Pro false (in i )=P1×Effect false (in i ) Pro true (in i )=P1×Effect true (in i ) Where P1 and P2 are user u i The proportion of fake topics and real topics spread among friends; Effect false and Effect true They are the topic influences under the influence of psychological benefits of related false topics and real topics in the historical topics participated by users; Finally, the calculation formula of evolutionary game theory is used, and the impact of user psychological benefits is taken into account to construct the final mutual influence between false and real topics.

10. The method for predicting the spread of false and real guiding topics based on user psychological benefits and adversarial generative networks according to claim 9 is characterized by: In step S5, PI-GCN is used as the prediction model, where the input of the prediction model is: Feature matrix X = N × A, where N represents the number of user nodes in the false topic propagation network, and A is the topic feature vector of each user node, which includes the user's own attribute characteristics and topic driving factor characteristics; Adjacency Matrix in Game Represents the relationship connection information between all users in the two topic spaces of false topics and real topics under the influence of user psychological benefits at time t; First, a two-layer graph convolutional neural network with an intermediate Dropout layer is used as a network rumor forwarding prediction model. First, the weights and biases are randomly initialized. Then, X is multiplied by W and the bias is added, and then multiplied by A^. After that, the RELU function is used as the activation function of this layer, and the Dropout operation is performed during model training. Finally, the SoftMax activation function is used to represent the convolution output as the probability value of different node classifications; it is expressed as: Z=f(X,A)=softmax(A^ReLU(A^XW 0 )W 1 ) Among them, W i is the weight matrix corresponding to the i-th layer network in the graph convolutional network, Let the model output Z = P(r, a, d|u i ), the specific definition is as follows: Among them, if the corresponding D = 1, then the potential user u is judged i The false topic will be forwarded in the next time period; if D = -1, the potential user u i will forward the real topic in the next time period; otherwise, potential user u i In the next period of time, we will not participate in the dissemination of false or true topics.

Citation Information

Patent Citations

  • Network rumor propagation prediction method based on user short-term emotion and evolutionary game

    CN115470991A

  • Propagation situation awareness prediction method based on multi-dimensional cognition and game theory

    CN117057944A

  • Tape measure with improved measurement convenience

    KR1020260061775A

  • Method for analysing social influence between internet forums and apparatus for the same

    KR102615014B1

  • Method and system to predict the likelihood of topics

    US20100280985A1