Sensitive topic prediction method based on user association strength and cognitive difference
By strengthening the processing of social network data and quantifying user emotions and cognitively, building a user emotional cognitive matrix and inputting a prediction network, the problems of early communication structure of sensitive topics in social networks are solved, and the accurate identification of sensitive topics and early suppression of communication are achieved.
Patent Information
- Application Number
- CN202510346482.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has problems of structural limitations, high user emotional intensity and cognitive differences in identifying and suppressing the early propagation of sensitive topics in social networks.
By acquiring and preprocessing user data, topic data and social network data, data enhancement processing is carried out to generate enhanced topics and social networks, quantify user emotional tendencies and cognitive influence, build a user emotional cognitive matrix, and input sensitive topic prediction networks for prediction.
Effectively identify and inhibit the spread of sensitive topics, overcome the scarcity of early communication structure, accurately represent the users' high emotional characteristics and cognitive differences in sensitive topics, and achieve the effect of inhibiting the spread of sensitive topics in the early stage.
Smart Images

Figure CN120217102A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of social network sensitive topic recognition, and particularly relates to a sensitive topic prediction method based on user association strength and cognitive difference. Background Art
[0002] In the current era when social media is flooded with a large amount of information, the spread of sensitive topics has become extremely complex. Sensitive topics are a collection of non-positive, negative, harmful, critical, and attacking public opinion information centered around a certain subject, including but not limited to enterprises, organizations, public figures, natural persons, etc. Sensitive topics often involve issues such as race, gender, religion, politics, and other highly controversial topics, and tend to easily trigger strong social reactions. Especially topics such as rumors and false information, these hot topics are often extremely harmful and difficult to distinguish between true and false. How to accurately detect sensitive topics and suppress their spread at an early stage has become a hot topic in academic research.
[0003] In recent years, the research on sensitive topic detection methods on social networks has received extensive attention from the academic community, and scholars have conducted a series of studies on it. According to different research focuses, these methods can be roughly divided into two categories: content-based and propagation feature-based. Content-based detection methods mainly detect by observing the change trends of text features and user features in the time axis during the topic propagation process. Detection methods based on the social propagation network structure mainly detect sensitive topics by capturing node features and propagation patterns in the network topology structure. However, the current sensitive topic recognition methods still have the following deficiencies:
[0004] 1. Problem of limited early propagation structure of sensitive topics. Due to the particularity of sensitive topics, in the early stage of propagation, the spread of sensitive topics is often restricted by social pressure, the government, and relevant laws, resulting in a relatively limited information dissemination path. At the same time, due to the incompleteness of information, users are skeptical about the topic, resulting in low user participation. These factors work together to cause the problem of limited early propagation structure of sensitive topics.
[0005] 2. Problem of high emotional intensity of user emotions. Sensitive topics can often quickly stimulate users' high emotional intensity reactions with obvious polarity and intensity. The ambiguity of information poses a cognitive challenge to users, and the amplification effect of social media platforms further exacerbates the spread and resonance of emotions. These factors are intertwined and jointly contribute to the high intensity problem of users' emotional reactions in the discussion of sensitive topics.
[0006] 3. Problem of differences in user perception. When a topic spreads, due to different preferences, users tend to have different tendencies in choosing information and activities. At the same time, with the existence of cognitive biases, such as the herd effect, authority bias, etc., these biases often lead to deviations in users' judgment of information. These factors jointly affect users' cognitive ways, resulting in different behavioral states of users when facing the same information. Summary of the Invention
[0007] To solve the above technical problems, the present invention proposes a sensitive topic prediction method based on user association strength and cognitive differences.
[0008] S1: Obtain user data, topic data, and social network data, and perform data preprocessing on them respectively.
[0009] S2: Perform data enhancement processing according to the topic data and social network data after data preprocessing to obtain enhanced topics and an enhanced social network, and obtain a user relationship matrix according to the enhanced social network.
[0010] S3: Quantify the emotional tendency of users towards topics according to user friend relationship data, user comment set data, and enhanced topics to obtain a user emotion quantification matrix Matrix.
[0011] S4: Quantify the cognitive influence of users on topics according to user basic information data and enhanced topics to obtain a user cognitive matrix PR.
[0012] S5: Concatenate the user emotion quantification matrix Matrix and the user cognitive matrix PR to obtain a user emotion and cognition matrix ECM.
[0013] S6: Input the user emotion and cognition matrix ECM and the user relationship matrix into a sensitive topic prediction network to obtain a prediction result.
[0014] Advantages of the present invention: First, the present invention collects user data, topic data, and social network data from social network platforms, performs data enhancement processing on the topic data and social network data to obtain enhanced topics and an enhanced social network, and then obtains a user relationship matrix, which accurately reflects the association strength between users in the enhanced social network and effectively overcomes the scarcity problem of the early social network structure. Second, the present invention quantifies the emotional tendency of users towards topics based on user friend relationship data, user comment set data, and enhanced topics to obtain a user emotion quantification matrix Matrix. Third, the present invention quantifies the cognitive influence of users on topics based on user basic information data and enhanced topics to obtain a user cognitive matrix PR. Finally, the user emotion quantification matrix Matrix and the user cognitive matrix PR are concatenated to obtain a user emotion and cognition matrix ECM. The user emotion and cognition matrix ECM can accurately represent the high-emotion characteristics of users towards sensitive topics and the differences in topic cognition among different users. Node2Vec is used to optimize the network structure and matrixize the emotional characteristics and cognitive characteristics of users. The user emotion and cognition matrix ECM and the user relationship matrix are input into a sensitive topic prediction network to obtain a prediction result. The present invention can effectively identify sensitive topics such as rumors and achieve the effect of suppressing the spread of sensitive topics in the early stage. Description of the Drawings
[0015] Figure 1 It is a schematic flowchart of an embodiment of the present invention;
[0016] Figure 2 It is a step flowchart of an embodiment of the present invention;
[0017] Figure 3 It is a schematic diagram of the process of data enhancement processing in an embodiment of the present invention;
[0018] Figure 4 It is a schematic diagram of the process of obtaining an enhanced social network based on enhanced topics in an embodiment of the present invention;
[0019] Figure 5 It is a schematic diagram of the process of quantifying the emotional tendency of users towards topics in an embodiment of the present invention;
[0020] Figure 6 It is a schematic diagram of the process of constructing a user emotion and cognition matrix in an embodiment of the present invention. Detailed Embodiments
[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0022] An embodiment of the present invention proposes a sensitive topic prediction method based on user association strength and cognitive difference. Referring to Figure 1 , Figure 2 as shown, the method includes:
[0023] S1: Obtain user data, topic data, and social network data, and perform data preprocessing on them respectively.
[0024] Specifically, collect the original data from the open APIs of each social network platform. The collected original data includes user data, topic data, and social network data. For example, the social network platforms include Weibo, WeChat, Twitter, etc. Among them, the user data includes: user basic information, such as the user's follows and followers, and the user's historical behavior data, such as likes, forwards, comments, collections, etc.; the topic data includes: target topics, topic sets (such as various topic keywords); the social network data includes the graph structure G of the social network, G=(V, E), where V represents the node set, the nodes represent users in the social network, and E represents the edge set, and the edges represent the connection relationship between users, or rather, the propagation relationship.
[0025] Usually, the obtained original data is unstructured and cannot be directly used for data analysis, and data preprocessing is required. Data preprocessing mainly includes data cleaning, and data cleaning can structure most of the unstructured data. The main specific methods of data cleaning include: deleting duplicate data, cleaning invalid nodes (such as some tourist data), etc.
[0026] S2: Perform data enhancement processing on the topic data and social network data after data preprocessing to obtain enhanced topics and an enhanced social network, and obtain a user relationship matrix according to the enhanced social network.
[0027] Figure 3 is a schematic diagram of the data enhancement processing process in the embodiment of the present invention. Figure 3 In, the data enhancement processing process includes: using the TF-IDF algorithm to find similar topics to the target topic according to the user comment set; mapping the propagation of the target topic and similar topics in the social network to a random walk in the social network graph G, extracting node (i.e., user) vectors, calculating the cosine similarity between nodes, finding similar nodes (or called shared nodes), and obtaining enhanced topics; obtaining an enhanced social network according to the enhanced topics, and further obtaining a user relationship matrix.
[0028] Referring to Figure 3 as shown, the specific process of the data enhancement processing includes:
[0029] S201: Use the TF-IDF algorithm to find similar topics to the target topic according to the user comment set.
[0030] It should be noted that the Term Frequency-Inverse Document Frequency (TF-IDF) algorithm is a commonly used statistical method in information retrieval and text mining, which is used to evaluate the importance of a word in a document or a set of documents (corpus). This algorithm consists of two parts: Term Frequency (TF) and Inverse Document Frequency (IDF).
[0031] Specifically, referring to Figure 3 As shown, obtain the target topic data and other potentially similar topic data from the user comment set. Use the TF-IDF algorithm to obtain the document vectors of the target topic and the document vectors of n other potentially similar topics respectively, and calculate the cosine similarity between the document vectors of the n other potentially similar topics and the document vector of the target topic; sort according to the cosine similarity, and select the top m similar topics similar to the target topic.
[0032] S202: Map the process of topic propagation in the social network to a random walk according to the social network graph structure G to generate a node sequence.
[0033] Specifically, use the Node2Vec algorithm to perform a random walk according to the graph structure of the social network. It should be noted that the Node2Vec algorithm is an algorithm for node representation learning in a network. It combines two graph traversal strategies: Depth-First Search (DFS) and Breadth-First Search (BFS), aiming to generate node embeddings that can effectively capture the characteristics of the network structure. The uniqueness of Node2Vec lies in its flexible random walk strategy, which allows simultaneous exploration of local and global network structures. Specifically, two parameters p and q are used to control the tendency of the random walk: the parameter p controls the tendency to return to the previous node (reverse movement). The larger the p value, the less likely it is to return to the previous node. The parameter q affects the tendency of the walk to move to deeper or shallower levels. The smaller the q value, the more inclined to explore areas far from the starting node; otherwise, it is more inclined to explore areas near the starting node.
[0034] Generally, the transition probability of the random walk is as follows:
[0035]
[0036] where P(c i= x|c i-1 = v) represents the probability of a node's random walk, c i = x represents the node (i.e., user) x in the social network graph G, c i-1 = v represents the node (i.e., user) v adjacent to node x in the social network graph G, π vx represents the non-normalized transition probability from node v to node x, generally defined as the weight of the edge, i.e., π vx = w vx , Z represents the normalization constant, f(v, x) ∈ E represents that the non-normalized nodes v and x belong to the edge set, and E represents the edge set in the graph.
[0037] To reduce the node information bias caused by indiscriminate random walks, p and q parameters are introduced to control the preference of random walks, while balancing depth-first search DFS and breadth-first search BFS, and improving the transition probability of random walks to obtain the node transition probability in the topic, which is specifically expressed as:
[0038] π vx = w vx ·α pq (t, x)
[0039]
[0040] In the formula, π vx represents the transition probability of the topic from the non-normalized node v to node x, w vx represents the weight of the edge G in the social network graph structure, α pq (t, x) represents the formula introduced for improving the random walk, d tx represents the distance between nodes t and x in the social network graph structure G, and both p and q parameters are adjustable parameters.
[0041] S203: Use the Skip-Gram algorithm to map the node sequence to node vectors.
[0042] To reduce the computational complexity, the generated node vectors are non-linearly reduced to two-dimensional vectors through the T-SEN algorithm.
[0043] It should be noted that the Skip-Gram algorithm is one of the two algorithms in the Word2Vec model, mainly used to generate word vectors in the field of natural language processing. The goal of Skip-Gram is to learn high-quality distributed word representations (i.e., word vectors) from a large amount of text corpus, and through these word vectors, the semantic relationships between words can be captured.
[0044] It should be noted that the T-SEN algorithm (t-distributed Stochastic Neighbor Embedding) is a non-linear algorithm for dimensionality reduction and visualization. Its core lies in mapping high-dimensional data points to a low-dimensional space (usually two-dimensional or three-dimensional) to preserve the local relationships between data points. The T-SEN algorithm measures the similarity of data points in the high-dimensional space, represents this similarity as a probability distribution, then constructs a similar probability distribution in the low-dimensional space, and finally minimizes the difference between the two distributions through the gradient descent method.
[0045] S204: Calculate the cosine similarity between the node vectors of the target topic and the similar topics, compare the cosine similarity with a threshold, and determine whether it is a shared node.
[0046] Specifically, set a threshold θ (θ is an adjustable parameter) to determine whether a node is a shared node isNode, which is specifically expressed as:
[0047]
[0048] In the formula, the value of isNode represents the degree of association between different user nodes. Both A and B represent user node vectors, and cos(A, B) represents the cosine similarity of the user node vectors. When cos(A, B) > θ, the two user nodes (A, B) are similar nodes, and the count of the shared node isNode is incremented by 1. When cos(A, B) < θ, the two user nodes (A, B) are not similar nodes, and the count of the shared node isNode is 1.
[0049] S205: Obtain the enhanced topic based on the shared nodes, obtain the enhanced social network based on the enhanced topic, and construct the user relationship matrix Matrix according to the enhanced social network value 。
[0050] Figure 4 This is a schematic diagram of the process of obtaining the enhanced social network based on the enhanced topic in the embodiments of the present invention. Figure 4 In it, Topic represents the topic, Node2Vec represents the Node2Vec algorithm, and Weight Union Graph represents the structure of the enhanced social network graph.
[0051] Specifically, refer to Figure 4As shown in the figure, select the top 2 most similar topics from the m similar topics, and find the shared nodes (i.e., similar users) between these two similar topics and the target similar topic according to the method of S204. After obtaining the shared nodes, merge the propagation relationship or forwarding relationship of this node in the similar topic into the target topic, and use the isNode value obtained in S204 as the weight of the edge therein, so as to obtain the enhanced social network. The enhanced social network represents the forwarding relationship graph of users for the topic.
[0052] Based on the correlation degree between users in the enhanced social network, construct the user relationship matrix Matrix value , which is specifically expressed as:
[0053] Matrix value =(isNode) m×n
[0054] In the formula, the value of the relationship matrix Matrix value represents the correlation strength value between users in the enhanced social network. isNode = 0 indicates that there is no direct forwarding relationship between nodes (i.e., users), that is, the relationship between nodes is a non-connected relationship. isNode ≠ 0 indicates that there is a direct forwarding relationship between nodes (i.e., users), that is, there is a connection relationship between nodes (i.e., users). m and n are the dimensions of the matrix.
[0055] After being processed by step S2 in the embodiment of the present invention, the problem of scarce early data caused by the particularity of sensitive topics is made up, the target data is further strengthened, and the generalization ability of the model and the detection accuracy of the model have a certain improvement effect.
[0056] Any user's perception of a topic is different and is affected by internal driving factors (i.e., the user's own emotions) and external driving factors (i.e., the emotions of the user's friends towards the topic). For example, if a user has a positive emotion towards a topic, they will actively like, forward, comment or collect on the social network and actively participate in the spread of the topic; if the user's friend has an opposite emotion (i.e., negative emotion) towards the topic, it will also affect the user's emotion towards this topic. The positive or negative emotion of the user towards the topic is in a game state and may change at any time. In the embodiment of the present invention, the topic emotion is quantified.
[0057] S3: Quantify the emotional tendency of users towards the topic according to the user friend relationship data, the user comment set data and the enhanced topic, and obtain the user emotion quantification matrix Matrix.
[0058] Figure 5 This is the schematic diagram of the quantification process of the emotional tendency of users towards the topic in the embodiment of the present invention.
[0059] The user's emotion is affected by various factors such as their own emotion and their friends' emotions. According to the internal emotional driving intensity Emo in (u i ,t i ) and the external emotional driving intensity Emo out (u i ,t i ), the multiple linear regression algorithm is used to analyze the user's complex emotions in a fine-grained manner.
[0060] Referring to Figure 5 shown, the quantification process of the user's emotional tendency towards the topic includes:
[0061] S301: Construct the emotional influence function I sup (u i ,t) for supporting the topic spread and the emotional influence function I opp (u i ,t) for opposing the topic spread, which are specifically expressed as:
[0062] I sup (u i ,t) = w0 + w1·log 10 Emo in-sup (u i ,t) + w2·log 10 Emo out-sup (u i ,t)
[0063] I opp (u i ,t) = w0 + w1·log 10 Emo in-opp (u i ,t) + w2·log 10 Emo out-opp (u i ,t)
[0064] In the formula, Emo in-sup represents the internal emotional driving factor for the user to support the topic spread, Emo in-opp represents the internal emotional driving factor for the user to oppose the topic spread, Emo out-sup represents the external emotional driving factors for the user's friends (also known as buddies) to support and oppose the topic spread, Emo out-opp represents the external emotional driving factor for the user's friends to oppose the topic spread, w0 represents the basic influence excluding the user's external and internal emotions, w1 represents the weight coefficient of the user's internal emotional driving factor, and w2 represents the weight coefficient of the user's external emotional driving factor. The logarithmic function log(.) is introduced to balance the internal and external emotional driving influences.
[0065] Generally speaking, there is often an obvious antagonistic and cooperative relationship between users' negative emotions and positive emotions. In order to more accurately quantify users' emotional tendencies, subjective game theory is used to quantify the mutual influence of users' emotions. At the same time, the sigmoid function is often used to describe the competition and antagonistic relationship between two opposing factors, and thus two emotional influences are defined.
[0066] S302: According to the emotional influence function I sup (u i , t) and I opp (u i , t), calculate the emotional influence of users supporting the topic I game-sup (u i , t) and the emotional influence of users opposing the topic I game-opp (u i , t) respectively, which are specifically expressed as:
[0067]
[0068] In the formula, I sup (u i , t) - I opp (u i , t) represents the factor of users' attitude towards supporting the topic, and I sup (u i , t) - I opp (u i , t) represents the factor of users' attitude towards opposing the topic. It is generally considered that the closer I game-sup (u i , t) and I game-opp (u i , t) are to 1, the greater the intensity of users' support or opposition to the topic spread.
[0069] S303: According to the magnitude of the two emotional forces I game-sup (u i , t) and I game-opp (u i , t), obtain the final emotional tendency Emo ui of users towards the topic, which is specifically expressed as:
[0070]
[0071] In the formula, represents that users support the topic, 1 represents the basic support degree, and emoIntensity sup represents the weighted value of the quantified support intensity; represents that users oppose the topic, 0 represents the basic opposition degree, and emoIntensityopp Indicates the weighted value of the opposition strength after quantization.
[0072] To better reflect the game between user emotions, the emotional weighted strength value emoIntensity att , which is specifically expressed as:
[0073]
[0074] In the formula, I game-sup (u i ,t)-I game-opp (u i ,t) represents the relative strength value of the user's support for the topic, and I game-opp (u i ,t)-I game-sup (u i ,t) represents the relative strength value of the user's opposition to the topic, and emoIntensity att ∈(0,1).
[0075] Based on the above, the precise quantization value of the user's emotion can be obtained, that is, the user emotion quantization matrix Matrix emo (topic), which is specifically expressed as:
[0076]
[0077] Compared with traditional sentiment analysis methods, such as only considering the user's sentiment from the text content of the user; the user sentiment quantization method based on subjective game theory and multiple linear regression in the embodiments of the present invention can more deeply consider the internal and external driving factors, emotional interaction, personality differences and dynamic changes of the user's sentiment. Through multi-dimensional analysis, this method provides a more accurate and comprehensive understanding of sentiment, which helps to better perform sentiment quantization.
[0078] Figure 6 It is a schematic diagram of the process of constructing the user emotion recognition matrix in the embodiments of the present invention.
[0079] S4: According to the user's basic information data and the enhanced topic, quantify the user's cognitive influence on the topic to obtain the user cognitive matrix PR.
[0080] Referring to Figure 6 shown, the specific process of calculating the user cognitive matrix PR includes:
[0081] S401: Calculate the basic cognitive vector of the user according to the user's basic information data.
[0082] Specifically, based on the user's basic information data, obtain the user's basic attributes, such as the number of followers (commonly known as the number of fans), the number of following, the number of likes, the number of forwards, the number of comments, and use the multiple linear regression algorithm to calculate the user's basic cognitive influence userCog(u i ), and then obtain the basic cognitive scores of all users in the social network, and obtain the basic cognitive vector C of all users, C = [userCog(u1), userCog(u2),..., userCog(u n )].
[0083] S402: Construct the cognitive weighted transfer matrix M' of the user based on the enhanced social network, and perform normalization processing on it to obtain the subsequent cognitive transfer matrix L.
[0084] Based on the enhanced social network, that is, the user's forwarding relationship graph, use the PageRank algorithm for iterative calculation.
[0085] Construct the user's original cognitive matrix M ij , which is specifically expressed as:
[0086]
[0087] Among them, OutDeg(vi) represents the out-degree of node v i (that is, user v i ).
[0088] Based on the enhanced social network and the user's basic cognitive vector C, after cognitive weighting, obtain the user's cognitive weighted transfer matrix M', which is specifically expressed as:
[0089] M' = M ⊙ (1 · C T )
[0090] Among them, ⊙ represents the Hadamard product, and C T represents the matrix transpose of the user's basic cognitive vector C.
[0091] Normalize the obtained cognitive weighted transfer matrix M' to obtain the normalized cognitive transfer matrix L, which is specifically expressed as:
[0092] L = diag(M' · 1) -1 · M'.
[0093] S403: Based on the user's basic cognitive vector C, use the PageRank algorithm to initialize the PR value of the user's cognitive vector, that is, initialize the PageRank vector, which is expressed as:
[0094]
[0095] PR (0) denotes the initialization of the PageRank vector, and ||C||1 denotes the L1 norm of the basic cognitive vector of the user, that is
[0096] S404: Iteratively calculate the PR value of the user's cognitive vector until the end condition is reached, and obtain the final user cognitive matrix PR.
[0097] Iteratively calculate the PR value of the user's cognitive vector, and its calculation formula is:
[0098]
[0099] In the formula, PR k+1 denotes the PR value of the user's cognitive vector calculated after the (k + 1)-th iteration, d denotes the damping factor, usually set to 0.85, and L T denotes the matrix transpose of L, and PR k denotes the PR value of the user's cognitive vector calculated after the k-th iteration.
[0100] When ||PR k+1 - PR k || < ε, stop the calculation. ε represents the set threshold for iteration, usually set to 0.001. Thus, the final user cognitive vector can be obtained as PR = [PR1, PR2,..., PR N . According to the calculated PR values of the user's cognitive vector under multiple factors, convert the user's cognitive vector into matrix form, that is, the user cognitive matrix PR.
[0101] In the embodiments of the present invention, the user's basic cognitive influence (number of followers, number of likes, number of forwards, number of comments, number of fans) is incorporated into the PageRank algorithm, which can not only more accurately quantify the user's cognitive influence, but also enhance the personalized analysis of the model and optimize the cognitive propagation path. This method can comprehensively measure the user's influence in a complex social network and provide strong support for more accurate cognitive analysis.
[0102] S5: Concatenate the user emotion quantification matrix Matrix and the user cognitive matrix PR to obtain the user emotion cognitive matrix ECM.
[0103] Refer to Figure 6As shown in the figure, specifically, in order to comprehensively consider the emotional factors of users, the user emotion matrix and the user's cognitive influence matrix are transformed into matrices in the same dimension for simple splicing to construct an Emotion Cognition Matrix (ECM). The embodiment of the present invention makes up for the limitation of previous studies that only detect sensitive topics from text content, integrates the user cognitive matrix PR into the user emotion quantification matrix Matrix, strengthens the features of the user's feature matrix, improves the prediction accuracy of the model, and at the same time enables the model to be more efficient and flexible, and better processes various types of data. S6: Input the user emotion cognition matrix ECM and the user relationship matrix into the sensitive topic prediction network to obtain a prediction result.
[0104] The sensitive topic prediction network includes a graph convolutional neural network GCN and a fully connected layer, and the fully connected layer performs binary classification.
[0105] It should be noted that the core idea of the Graph Convolutional Network (GCN) is to update the representation of the current node by aggregating the information of neighbor nodes. GCN usually uses a multi-layer neural network to train node embeddings. Starting from the adjacency matrix and node feature matrix of the graph, it learns node embeddings by propagating and aggregating the information of neighbor nodes layer by layer. Each layer of GCN can be regarded as a process of message passing, where nodes update the representation of the nodes by aggregating the information of neighbor nodes.
[0106] The output of the graph convolutional neural network is expressed as:
[0107] p(topic) = softmax{GCN(ECM(matrix) + Matrix(value)}
[0108] p(topic) represents the class probability output by the graph convolutional neural network, ECM(matrix) represents the user emotion cognition matrix, Matrix(value) represents the user relationship matrix, and GCN(.) represents the graph convolutional neural network model.
[0109] Input the prediction probability value p(topic) of the graph convolutional neural network into the fully connected layer for binary classification. Set the threshold to θ. When the prediction probability value p(topic) is greater than θ, it is determined that the topic is a sensitive topic, otherwise it is not a sensitive topic.
[0110] The output result of the fully connected layer, that is, the classification result isTopic, is specifically expressed as:
[0111]
[0112] Among them, isTopic represents the topic type. isTopic = 1 indicates that the topic is a sensitive topic; isTopic = 0 indicates that the topic is a non-sensitive topic, and θ is an adjustable parameter.
[0113] In the network training stage of the sensitive topic prediction network, its training process is as described in the above steps S1 - S6, only the data collected in S1 is different. In the network training stage, the data collected in S1 are all historical data. For the same steps, they will not be elaborated here.
[0114] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium can include: ROM, RAM, magnetic disk, optical disc, etc.
[0115] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A sensitive topic prediction method based on user association strength and cognitive differences, characterized in that: include: Obtain user data, topic data, and social network data, and perform data preprocessing on them respectively; Perform data enhancement processing on the topic data and social network data after data preprocessing to obtain enhanced topics and enhanced social networks, and obtain a user relationship matrix based on the enhanced social networks; According to the user's friend relationship data, user comment collection data and enhanced topics, the user's emotional tendency towards the topic is quantified to obtain the user's emotional quantification matrix Matrix; According to the user's basic information data and enhanced topics, the user's cognitive influence on the topic is quantified to obtain the user cognitive matrix PR; The user emotion quantization matrix Matrix and the user cognitive matrix PR are concatenated to obtain the user emotion cognitive matrix ECM; The user's emotion cognition matrix ECM and user relationship matrix are input into a sensitive topic prediction network to obtain a prediction result.
2. The sensitive topic prediction method based on user association strength and cognitive differences according to claim 1 is characterized in that: User data includes basic user information and historical behavior data of users, topic data includes target topics and topic sets, and social network data includes the graph structure G of the social network.
3. The sensitive topic prediction method based on user association strength and cognitive differences according to claim 1 is characterized in that: The specific process of the data enhancement processing includes: Using the TF-IDF algorithm, we find similar topics to the target topic based on the user comment set; Map the topic propagation process in the social network into a random walk according to the social network graph structure G, generate a node sequence, and map the node sequence into a node vector; Calculate the cosine similarity between the node vectors in the target topic and the similar topic, compare the cosine similarity with a threshold, and determine whether it is a shared node; Based on the shared nodes, we get the enhanced topics, based on the enhanced topics, we get the enhanced social network, and based on the enhanced social network, we build the user relationship matrix Matrix value .
4. The sensitive topic prediction method based on user association strength and cognitive differences according to claim 1, characterized in that: The process of quantifying users' sentiment towards topics includes: Constructing the emotional influence function of user support topic propagation I sup (u i ,t) and the emotional influence function I of the opposing topic propagation opp (u i ,t); According to the emotional influence function I sup (u i ,t) and I opp (u i , t), respectively calculate the user support topic sentiment influence I game-sup (u i ,t) and the emotional influence of opposing topics I game-opp (u i ,t); According to the emotional influence I game-sup (u i ,t) and I game-opp (u i ,t), and obtain the user’s final emotional tendency Emo ui .
5. The sensitive topic prediction method based on user association strength and cognitive differences according to claim 1, characterized in that: The specific process of calculating the user recognition matrix PR includes: According to the user's basic information data, the user's basic cognitive vector is calculated; According to the enhanced social network, the user's cognitive weighted transfer matrix M' is constructed and normalized to obtain the cognitive transfer matrix L; Based on the user's basic cognitive vector C, use the PageRank algorithm to initialize the user's cognitive vector PR value; The user's cognitive vector PR value is iteratively calculated until the end condition is reached, and the final user cognitive matrix PR is obtained.
6. The sensitive topic prediction method based on user association strength and cognitive differences according to claim 1, characterized in that: The sensitive topic prediction network includes a graph convolutional neural network (GCN) and a fully connected layer.