Social network hot event identification method based on graph neural network
Patent Information
- Application Number
- CN202610954120.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]首先,基于关键词匹配和统计分析的方法严重依赖预设规则和人工设定的参数,难以适应社交网络中语言表达的多样性,社交媒体用户的表达方式往往带有非正式语句、网络流行语、情感倾向因素,仅依靠关键词匹配难以捕捉热点事件的真实语义容易导致误报或漏报,此外,统计分析方法依赖于时间窗口内的频率波动,无法识别信息在网络拓扑结构中的传播模式,难以发现真正具有影响力的热点事件
[0078](1)本发明通过引入自适应信息熵优化机制,针对热点事件传播中的关键信息节点进行动态加权使得模型能够自动识别高价值传播节点,提高热点事件识别的精准度,通过计算初始信息熵、信息多样性熵和信息流动熵综合衡量节点的信息传播能力,并自适应调整邻居特征聚合权重,有效提升了热点事件传播路径的建模精度,使得模型能够捕捉复杂的事件扩散模式,从而降低误报率和漏报率。
Smart Images

Figure CN122817564A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of social network hot topic technology, and in particular to a method for identifying social network hot topic events based on graph neural networks. Background Technology
[0002] With the rapid development of social networks, social media has become an important channel for information dissemination. Hot topics are characterized by their suddenness, rapid spread, and wide impact. However, due to the large user base and diverse data types of social networks, how to accurately identify hot topics from massive amounts of social data has become an important research direction in the field of social network analysis.
[0003] Currently, the identification of trending events on social networks mainly relies on keyword matching, statistical analysis, and traditional machine learning models. Keyword matching methods detect high-frequency words or co-occurrence patterns in social networks using pre-defined keyword or topic models to identify potential trending events. Statistical analysis methods typically employ TF-IDF, word frequency change analysis, and time series fluctuation detection to statistically model social network text data to discover anomalous surges in information. Traditional machine learning methods utilize support vector machines and random forest classification models to extract features from social media data and train classifiers to automatically detect the occurrence of trending events. Although existing methods can identify trending events on social networks to some extent, they still have many limitations.
[0004] First, keyword matching and statistical analysis methods rely heavily on preset rules and manually set parameters, making it difficult to adapt to the diversity of language expression in social networks. Social media users often use informal language, internet slang, and emotional biases in their expressions. Relying solely on keyword matching makes it difficult to capture the true semantics of trending events, which can easily lead to false positives or false negatives. In addition, statistical analysis methods rely on frequency fluctuations within a time window and cannot identify the propagation patterns of information in the network topology, making it difficult to discover truly influential trending events.
[0005] Secondly, traditional machine learning methods mainly rely on text features for classification, ignoring the topological structure of social networks. In social networks, the spread of trending events depends not only on the text content but also on the social relationships, interaction behaviors, and information flow patterns among users. The same content may spread more widely if posted by a high-influence user. Traditional methods struggle to effectively model user influence and dissemination paths. Furthermore, traditional machine learning methods typically require manually designed features, rely on a large amount of labeled data during training, have limited generalization ability, and are difficult to adapt to the dynamic changes in the social network environment.
[0006] In summary, existing technologies have significant shortcomings in terms of accuracy in identifying trending events, utilization of social network structural information, and computational efficiency, making it difficult to meet the demand for efficient and accurate identification of trending events in a social network environment. Therefore, there is an urgent need for a new method that can fully combine the topological structure and content information of social networks to improve the accuracy of trending event identification while taking into account computational efficiency, so as to meet the real-time trending event detection needs in a large-scale data environment of social networks. Summary of the Invention
[0007] One objective of this invention is to propose a method for identifying trending events on social networks based on graph neural networks. This invention overcomes the limitations of existing technologies for identifying trending events on social networks and improves the accuracy, feature learning ability, and computational efficiency of trending event detection.
[0008] A method for identifying trending events on social networks based on graph neural networks according to an embodiment of the present invention includes the following steps:
[0009] S1. Construct a social network graph based on social network data. The social network graph includes nodes and edges, where the nodes represent users or information units in the social network data, and the edges represent the interactive and propagation relationships between users, such as following, forwarding, and commenting.
[0010] S2. Calculate the information entropy of each node in the social network graph using the node feature representation, generate the information entropy result of each node, and measure the information value of each node in the social network data;
[0011] S3. Calculate the adaptive information entropy weight matrix for each node based on the node information entropy results;
[0012] S4. Use graph neural networks to aggregate node features in a social network graph. During the aggregation process, the graph neural network uses an adaptive information entropy weight matrix to effectively aggregate the features of adjacent nodes to the target node and update the feature representation of the target node.
[0013] S5. Based on the updated node feature representation, the graph neural network is trained using the variational information entropy optimization loss function, so that the information entropy of hot event propagation nodes tends to be maximized and the information entropy of ordinary nodes tends to be minimized, thereby obtaining the optimized graph neural network model.
[0014] S6. The optimized graph neural network model is applied to the real-time social network data stream to construct a dynamic social network graph from the real-time social network data, and social network hot events are identified based on the node feature aggregation results and the adaptive information entropy weight matrix.
[0015] S7. Output the results of social network hot topic identification, including the topic of the hot topic, key propagation nodes and propagation path of the event.
[0016] Optionally, S1 includes the following steps:
[0017] S11. Collect social network data, wherein the social network data includes user interaction data, content data, and timestamp information indicating the time when user behavior occurs, wherein the user interaction data includes follow relationships, forwarding behavior, and comment interaction between users, and wherein the content data includes text, image, and video information posted by users;
[0018] S12. Analyze social network data and construct a social network graph based on user interaction data and timestamp information. The social network graph is represented as a directed weighted graph:
[0019] G = (V, E, W);
[0020] Where V is the set of nodes in the social network, and the nodes in the node set are... Let E be a user or information unit in the social network data, and let E be the set of edges in the social network. For users With users The interactions between them, including following, forwarding, and commenting, Let the set of edge weights in the social network graph be the weights in the set of edge weights. The weights, calculated based on user interaction frequency, interaction content similarity, or time decay factor, are used to measure the strength of the propagation relationship between nodes.
[0021] ;
[0022] in, Indicates user With users In social networks, the frequency of interaction matters; the higher the frequency, the greater the weight. Indicates user With users The similarity between the content is calculated using cosine similarity. The time decay function represents the time of interaction. Regarding the impact on edge weights, earlier interactions are assigned lower weights. , , This is a hyperparameter.
[0023] Optionally, S2 includes the following steps:
[0024] S21. Extract feature representations for each node in the social network graph, including user content features, user interaction features, and time-series features. The user content features are extracted based on text, image, or video information posted by users, and the user interaction features are based on edge weights. The extracted time series features are used to characterize the evolution pattern of user behavior over time.
[0025] S22. Calculate the initial information entropy of each node in the social network graph, wherein the initial information entropy is used to measure the information distribution characteristics of the node in the social network:
[0026] ;
[0027] in, For nodes The initial information entropy, For nodes The set of neighboring nodes, For nodes Propagate information to neighboring nodes The probability of propagation;
[0028] S23. Calculate the information diversity entropy of the node, wherein the information diversity entropy is used to characterize the category distribution of the information received by the node:
[0029] ;
[0030] in, For nodes The information diversity entropy, where C is the set of all information categories. For nodes The probability of receiving information of category c;
[0031] S24. Calculate the information flow entropy of the nodes, which measures the uniformity of information propagation among nodes in the social network:
[0032] ;
[0033] in, For nodes Information flow entropy For nodes The total propagation weight;
[0034] S25. Calculate the node information entropy results. Combine the node information entropy results with the initial information entropy, information diversity entropy, and information flow entropy to characterize the role of nodes in the propagation of hot events:
[0035] ;
[0036] in, For nodes The comprehensive information entropy, , , These are the weight parameters.
[0037] Optionally, S3 includes the following steps:
[0038] S31. Read the node information entropy result The node information entropy result is used to measure the information value of each node in the social network graph during the propagation of hot events. The larger the node information entropy result, the more significant the node's role in the propagation of hot events.
[0039] S32. Calculate the adaptive information entropy normalization weights for each node in the social network graph. These normalization weights are used to map information entropy values at different scales to a unified interval, resulting in the normalized node information entropy. ;
[0040] S33. Calculate the adaptive information entropy weights of the nodes. These weights are used to adjust the aggregation strength of neighbor features in the graph neural network.
[0041] ;
[0042] in, For nodes Adaptive information entropy weights, For adaptive adjustment coefficient, Use the Sigmoid activation function;
[0043] S34. Calculate the feature aggregation weights of each node in the social network graph relative to its neighboring nodes. The feature aggregation weights are calculated by combining the adaptive information entropy weights and the edge weights of the neighboring nodes:
[0044] ;
[0045] in, For nodes For neighboring nodes Weights for feature aggregation;
[0046] S35. Calculate the adaptive information entropy weight matrix that ultimately guides the aggregation of neighbor features in the graph neural network. :
[0047] .
[0048] Optionally, S4 includes the following steps:
[0049] S41. Read the feature representations of each node in the social network graph at the l-th layer of the graph neural network. ,in Represents a node In the feature vector of the l-th layer of the graph neural network, the feature vector captures the nodes. The basic attributes of trending events on social networks;
[0050] S42. Using an adaptive information entropy weight matrix For nodes Aggregate the features of neighboring nodes:
[0051] ;
[0052] in, This is the information entropy modulation coefficient, used to amplify or attenuate the effect of normalized comprehensive information entropy on feature aggregation. As a parameter sensitive to feature similarity, Represents a node with neighboring nodes The squared Euclidean distance between the feature representations in the l-th layer. These are the residual connectivity coefficients, used to preserve nodes. In the feature information of the previous layer, This is the normalization constant.
[0053] Optionally, S5 includes the following steps:
[0054] S51. Based on the updated node feature representation The predicted information entropy of a node is calculated, and this entropy is used to measure the importance of a node in the propagation of hot events.
[0055] ;
[0056] in, For nodes Predictive information entropy, For nodes Propagate information to neighboring nodes Predicted propagation probability:
[0057] ;
[0058] in, For nodes and nodes Feature similarity at layer l+1 of the graph neural network:
[0059] ;
[0060] S52. Calculate variational information entropy to optimize the loss function, so that the information entropy of hot event propagation nodes tends to be maximized, and the information entropy of ordinary nodes tends to be minimized:
[0061] ;
[0062] in, To optimize the loss for variational information entropy, For nodes The true label is 1 for nodes that spread hot events, and 0 for ordinary nodes;
[0063] S53. Gradient descent is used to optimize the graph neural network model. The parameters of the graph neural network model are updated by calculating the gradient of the loss function based on the variational information entropy, so that the graph neural network model is continuously optimized during training, thereby improving the performance of hotspot event recognition.
[0064] ;
[0065] in, Let be the parameters of the graph neural network model at the t-th iteration. For learning rate, The gradient of the loss function with respect to network parameters is optimized for variational information entropy.
[0066] Optionally, S7 includes the following steps:
[0067] S71. Predictive Information Entropy Based on Hotspot Event Propagation Nodes Identify the key dissemination nodes of trending events:
[0068] ;
[0069] Where K is the set of key propagation nodes, For nodes Predictive information entropy, For nodes with neighboring nodes Edge weights between them The information entropy threshold, These are the propagation intensity thresholds; only nodes that meet these two threshold conditions are identified as key nodes in the propagation of hotspot events.
[0070] S73. Calculate the propagation path of hot events, wherein the propagation path is constructed based on the information flow between key propagation nodes:
[0071] ;
[0072] Where P represents the optimal propagation path of the trending event. Given the set of all possible propagation paths on the key propagation node set K, the optimal solution for each propagation path is based on the propagation weights of all edges along the path. and target node Predictive information entropy The weighted sum and maximization calculation is obtained;
[0073] S74. Determine the event theme of the hot topic, which is calculated based on the content characteristics of the hot topic's propagation nodes:
[0074] ;
[0075] Where T represents the topic category of the trending event, and C represents the set of all topic categories. As a key propagation node The probability of belonging to topic category c is determined by maximizing the sum of the probabilities of propagation nodes in a certain category;
[0076] S75. Output the identification results of social network hot events, the identification results including the event topic T, the set of key propagation nodes K, and the event propagation path P.
[0077] The beneficial effects of this invention are:
[0078] (1) This invention introduces an adaptive information entropy optimization mechanism to dynamically weight key information nodes in the propagation of hot events, enabling the model to automatically identify high-value propagation nodes and improve the accuracy of hot event identification. By calculating the initial information entropy, information diversity entropy and information flow entropy to comprehensively measure the information propagation capability of nodes, and adaptively adjusting the neighbor feature aggregation weight, the modeling accuracy of hot event propagation path is effectively improved, enabling the model to capture complex event diffusion patterns, thereby reducing false alarm rate and false negative rate.
[0079] (2) This invention proposes an adaptive information entropy-driven feature aggregation method. By constructing an adaptive information entropy weight matrix, the graph neural network can prioritize the aggregation of node features with high propagation influence. It also dynamically adjusts the feature aggregation method in conjunction with the propagation structure of hot events. When calculating node features, the information entropy modulation coefficient and feature similarity weighting strategy are adopted, which effectively strengthens the role of high-influence nodes in the propagation of hot events and suppresses the noise interference of low-influence nodes.
[0080] (3) The present invention proposes a variational information entropy optimization loss function that maximizes the information entropy of hot event propagation nodes and minimizes the information entropy of ordinary nodes, thereby optimizing the training objective of the hot event detection model. During the training process, a gradient descent optimization strategy is used in combination with the predicted information entropy for backpropagation, so that the high-influence nodes of hot events can learn more features during the training process, while the influence of non-hot event nodes is effectively suppressed. In addition, the loss function design of the present invention reduces the computational cost of the model in large-scale social network data and improves the training efficiency. Attached Figure Description
[0081] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0082] Figure 1 This is a flowchart of a method for identifying hot topics in social networks based on graph neural networks, as proposed in this invention. Detailed Implementation
[0083] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0084] refer to Figure 1 A method for identifying trending events on social networks based on graph neural networks includes the following steps:
[0085] S1. Construct a social network graph based on social network data. The social network graph includes nodes and edges, where nodes represent users or information units in the social network data, and edges represent the interactive and propagation relationships between users, such as following, forwarding, and commenting.
[0086] S2. Calculate the information entropy of each node in the social network graph using node feature representation, generate the information entropy results of each node, and measure the information value of each node in the social network data.
[0087] S3. Calculate the adaptive information entropy weight matrix for each node based on the node information entropy results;
[0088] S4. Use graph neural networks to aggregate node features in a social network graph. During the aggregation process, the graph neural network uses an adaptive information entropy weight matrix to effectively aggregate the features of adjacent nodes to the target node and update the feature representation of the target node.
[0089] S5. Based on the updated node feature representation, the graph neural network is trained using the variational information entropy optimization loss function, so that the information entropy of hot event propagation nodes tends to be maximized and the information entropy of ordinary nodes tends to be minimized, thereby obtaining the optimized graph neural network model.
[0090] S6. The optimized graph neural network model is applied to the real-time social network data stream to construct a dynamic social network graph from the real-time social network data, and social network hot events are identified based on the node feature aggregation results and the adaptive information entropy weight matrix.
[0091] S7. Output the results of social network hot topic identification, including the topic of the hot topic, key propagation nodes and propagation path of the event.
[0092] In this embodiment, S1 includes the following steps:
[0093] S11. Collect social network data, which includes user interaction data, content data, and timestamp information indicating when user behavior occurs. User interaction data includes follow relationships, forwarding behavior, and comment interactions between users. Content data includes text, images, and video information posted by users.
[0094] S12. Analyze social network data and construct a social network graph based on user interaction data and timestamp information. The social network graph is represented as a directed weighted graph:
[0095] G = (V, E, W);
[0096] Where V is the set of nodes in the social network, and the nodes in the node set are... Let E be a user or information unit in the social network data, and let E be the set of edges in the social network. For users With users The interactions between them, including following, forwarding, and commenting, Let the set of edge weights in the social network graph be the weights in the set of edge weights. The weights are calculated based on user interaction frequency, similarity of interaction content, or time decay factor, and are used to measure the strength of the propagation relationship between nodes.
[0097] ;
[0098] in, Indicates user With users In social networks, the frequency of interaction matters; the higher the frequency, the greater the weight. Indicates user With users The similarity between the content is calculated using cosine similarity. The time decay function represents the time of interaction. Regarding the impact on edge weights, earlier interactions are assigned lower weights. , , This is a hyperparameter.
[0099] In this embodiment, S2 includes the following steps:
[0100] S21. Extract feature representations for each node in the social network graph, including user content features, user interaction features, and time-series features. User content features are extracted based on text, image, or video information posted by users, while user interaction features are based on edge weights. The extracted time-series features are used to characterize the evolution of user behavior over time.
[0101] S22. Calculate the initial information entropy for each node in the social network graph. The initial information entropy is used to measure the information distribution characteristics of the node in the social network:
[0102] ;
[0103] in, For nodes The initial information entropy, For nodes The set of neighboring nodes, For nodes Propagate information to neighboring nodes The probability of propagation;
[0104] S23. Calculate the information diversity entropy of the node. The information diversity entropy is used to characterize the distribution of categories of information received by the node:
[0105] ;
[0106] in, For nodes The information diversity entropy, where C is the set of all information categories. For nodes The probability of receiving information of category c;
[0107] S24. Calculate the information flow entropy of nodes. Information flow entropy is used to measure the uniformity of information propagation among nodes in a social network.
[0108] ;
[0109] in, For nodes Information flow entropy For nodes The total propagation weight;
[0110] S25. Calculate the node information entropy results. Combine the node information entropy results with the initial information entropy, information diversity entropy, and information flow entropy to characterize the role of nodes in the propagation of hot events:
[0111] ;
[0112] in, For nodes The comprehensive information entropy, , , These are the weight parameters.
[0113] In this embodiment, S3 includes the following steps:
[0114] S31. Read the node information entropy result The node information entropy result is used to measure the information value of each node in the social network graph during the propagation of hot events. The larger the node information entropy result, the more significant the node's role in the propagation of hot events.
[0115] S32. Calculate the adaptive information entropy normalization weights for each node in the social network graph. The normalization weights are used to map information entropy values at different scales to a unified interval, thus obtaining the normalized node information entropy. ;
[0116] S33. Calculate the adaptive information entropy weights of the nodes. These weights are used to adjust the aggregation strength of neighbor features in the graph neural network.
[0117] ;
[0118] in, For nodes Adaptive information entropy weights, For adaptive adjustment coefficient, Use the Sigmoid activation function;
[0119] S34. Calculate the feature aggregation weights of each node in the social network graph relative to its neighboring nodes. The feature aggregation weights are calculated by combining the adaptive information entropy weights and the edge weights of the neighboring nodes:
[0120] ;
[0121] in, For nodes For neighboring nodes Weights for feature aggregation;
[0122] S35. Calculate the adaptive information entropy weight matrix that ultimately guides the aggregation of neighbor features in the graph neural network. :
[0123] .
[0124] In this embodiment, S4 includes the following steps:
[0125] S41. Read the feature representations of each node in the social network graph at the l-th layer of the graph neural network. ,in Represents a node In the feature vector of the l-th layer of the graph neural network, the feature vector captures the nodes. The basic attributes of trending events on social networks;
[0126] S42. Using an adaptive information entropy weight matrix For nodes Aggregate the features of neighboring nodes:
[0127] ;
[0128] in, This is the information entropy modulation coefficient, used to amplify or attenuate the effect of normalized comprehensive information entropy on feature aggregation. As a parameter sensitive to feature similarity, Represents a node with neighboring nodes The squared Euclidean distance between the feature representations in the l-th layer. These are the residual connectivity coefficients, used to preserve nodes. In the feature information of the previous layer, This is the normalization constant.
[0129] In this embodiment, S5 includes the following steps:
[0130] S51. Based on the updated node feature representation The predicted information entropy of a node is calculated, and this entropy is used to measure the importance of a node in the propagation of hot events.
[0131] ;
[0132] in, For nodes Predictive information entropy, For nodes Propagate information to neighboring nodes Predicted propagation probability:
[0133] ;
[0134] in, For nodes and nodes Feature similarity at layer l+1 of the graph neural network:
[0135] ;
[0136] S52. Calculate variational information entropy to optimize the loss function, so that the information entropy of hot event propagation nodes tends to be maximized, and the information entropy of ordinary nodes tends to be minimized:
[0137] ;
[0138] in, To optimize the loss for variational information entropy, For nodes The true label is 1 for nodes that spread hot events, and 0 for ordinary nodes;
[0139] S53. Gradient descent is used to optimize the graph neural network model. The parameters of the graph neural network model are updated by calculating the gradient of the loss function based on the variational information entropy, so that the graph neural network model is continuously optimized during training, thereby improving the performance of hotspot event recognition.
[0140] ;
[0141] in, Let be the parameters of the graph neural network model at the t-th iteration. For learning rate, The gradient of the loss function with respect to network parameters is optimized for variational information entropy.
[0142] In this embodiment, S7 includes the following steps:
[0143] S71. Predictive Information Entropy Based on Hotspot Event Propagation Nodes Identify the key dissemination nodes of trending events:
[0144] ;
[0145] Where K is the set of key propagation nodes, For nodes Predictive information entropy, For nodes with neighboring nodes Edge weights between them The information entropy threshold, These are the propagation intensity thresholds; only nodes that meet these two threshold conditions are identified as key nodes in the propagation of hotspot events.
[0146] S73. Calculate the propagation path of hot events, which is constructed based on the information flow between key propagation nodes:
[0147] ;
[0148] Where P represents the optimal propagation path of the trending event. Given the set of all possible propagation paths on the key propagation node set K, the optimal solution for each propagation path is based on the propagation weights of all edges along the path. and target node Predictive information entropy The weighted sum and maximization calculation is obtained;
[0149] S74. Determine the event theme of the hot topic. The event theme is calculated based on the content characteristics of the hot topic's propagation nodes:
[0150] ;
[0151] Where T represents the topic category of the trending event, and C represents the set of all topic categories. As a key propagation node The probability of belonging to topic category c is determined by maximizing the sum of the probabilities of propagation nodes in a certain category;
[0152] S75. Output the identification results of trending events on social networks. The identification results include the event topic T, the set of key propagation nodes K, and the event propagation path P.
[0153] Example 1:
[0154] In June 2024, on a social media platform, user "@A" posted a tweet that read, "A certain brand of smart car suddenly accelerated and the brakes failed! A near crash! Is autonomous driving really reliable?" The tweet included an 18-second video showing the vehicle suddenly accelerating, the driver exclaiming in alarm, and then manually taking over the steering wheel a few seconds later. Within 10 minutes of posting, the tweet received 5,000 likes, 2,000 retweets, and 3,000 comments. As time went on, the number of retweets rapidly increased, surpassing 50,000 retweets within 20 minutes. Several well-known automotive bloggers and media accounts joined the discussion. One prominent tech blogger, "@B," retweeted the post 30 minutes later and commented, "The safety of autonomous driving systems is once again being questioned. Should car manufacturers bear responsibility?" This retweet received 12,000 likes and 5,000 retweets within 15 minutes, quickly escalating the incident.
[0155] After an event occurs, the method of this invention immediately initiates a hotspot event identification process to monitor the spread of the event in real time and compares it with traditional methods.
[0156] Within the first hour following the incident, this invention crawled 1,200,000 relevant posts, 5,000,000 comments, and 8,500,000 reposts from social media platforms in real time, and constructed a social network graph containing 620,000 user nodes and 4,200,000 edges, with edge weights calculated based on user interaction intensity. In contrast, traditional TF-IDF statistical methods could only extract 350,000 posts containing keywords such as "smart car," "brake failure," and "autonomous driving," missing a large amount of relevant information expressed in non-standard ways.
[0157] The method of this invention calculates the information entropy of all user nodes and identifies high-influence dissemination nodes, including the tech blogger "@B" (information entropy value 0.89), the automotive forum "@C" (information entropy value 0.85), and the news media "@D" (information entropy value 0.83). Among them, "@B" generated over 100,000 discussions within 30 minutes of the event's outbreak and was identified as the secondary dissemination center of the event. In contrast, traditional methods based on TF-IDF statistics failed to identify the dissemination influence of this node, resulting in incomplete identification of the hot topic's propagation path.
[0158] Within two hours, this invention calculated the main propagation path of the trending event and visualized its propagation network. The event was initially initiated by "@A", spread to the technology community via "@B", reached the automotive forum "@C" within one hour, and finally spread to the news media "@D", forming a cross-industry discussion. In comparison, the propagation path identification of traditional GNN methods only covered 68.2% of the actual propagation paths, while the method of this invention achieved an accuracy rate of 95.6%.
[0159] Within three hours of the incident, the event topic classification model of this invention analyzed all relevant posts and found that the discussion mainly focused on the following three aspects:
[0160] 1. Safety of autonomous driving (47.2%): Users are concerned about the reliability of autonomous driving systems and similar accident cases in the past.
[0161] 2. Responsibility of automakers (31.8%): Discussion on whether automakers should be held responsible for system failures and issues related to consumer rights protection.
[0162] 3. Rumors and misleading information (21.0%): Some users posted unverified information, such as "the car owner had modified the braking system" or "the driver did not use the automatic driving mode".
[0163] In contrast, topic analysis based on the TF-IDF statistical method can only identify the two keywords "autonomous driving" and "brake failure," and cannot accurately distinguish different discussion topics, resulting in coarse event classification and making it difficult to conduct accurate public opinion analysis.
[0164] The method of this invention is trained on a GPU server and its computational efficiency is compared with that of traditional methods. The method of this invention improves the computational efficiency by 48.0% compared with traditional GNN and by 66.7% compared with TF-IDF method, and reduces the demand for computing resources by 24.0% when processing large-scale data.
[0165] After event detection is completed, the false alarm rate and accuracy of the method of the present invention are compared with those of other methods. The false alarm rate of the present invention is reduced by 76.8% compared with the keyword matching method and by 66.7% compared with the TF-IDF method, while the accuracy of hot event identification is improved by 18.1%.
[0166] This embodiment demonstrates the significant advantages of the method of the present invention in terms of accurate identification of hot events, reconstruction of propagation paths, computational efficiency, and control of false alarm rate through real hot events on social media. The method of the present invention can not only efficiently and in real time detect hot events on social networks, but also accurately identify key propagation nodes and reconstruct complete propagation paths, significantly improving the accuracy and reliability of social network public opinion analysis. Experimental data shows that the detection accuracy, computational efficiency, and resource consumption of the present method are greatly optimized compared with traditional methods, and it is suitable for application scenarios such as emergency detection, rumor tracing, and public opinion analysis in large-scale social network environments.
[0167] This invention introduces an adaptive information entropy optimization mechanism to dynamically weight key information nodes in the propagation of hot events, enabling the model to automatically identify high-value propagation nodes and improve the accuracy of hot event identification. By calculating the initial information entropy, information diversity entropy, and information flow entropy to comprehensively measure the information propagation capability of nodes, and adaptively adjusting the neighbor feature aggregation weights, the model effectively improves the modeling accuracy of hot event propagation paths, enabling the model to capture complex event diffusion patterns, thereby reducing false positive and false negative rates.
[0168] This invention proposes an adaptive information entropy-driven feature aggregation method. By constructing an adaptive information entropy weight matrix, the graph neural network can prioritize the aggregation of node features with high propagation influence. The method is dynamically adjusted in conjunction with the propagation structure of hot events. When calculating node features, an information entropy modulation coefficient and feature similarity weighting strategy are adopted, which effectively enhances the role of high-influence nodes in the propagation of hot events and suppresses the noise interference of low-influence nodes.
[0169] This invention proposes a variational information entropy optimization loss function that maximizes the information entropy of hot event propagation nodes and minimizes the information entropy of ordinary nodes, thereby optimizing the training objective of the hot event detection model. During training, a gradient descent optimization strategy is used in conjunction with predicted information entropy for backpropagation, allowing high-influence nodes of hot events to learn more features during training, while effectively suppressing the influence of non-hot event nodes. In addition, the loss function design of this invention reduces the computational cost of the model in large-scale social network data and improves training efficiency.
[0170] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for identifying trending events on social networks based on graph neural networks, characterized in that, Includes the following steps: S1. Construct a social network graph based on social network data. The social network graph includes nodes and edges, where the nodes represent users or information units in the social network data, and the edges represent the interactive and propagation relationships between users, such as following, forwarding, and commenting. S2. Calculate the information entropy of each node in the social network graph using the node feature representation, generate the information entropy result of each node, and measure the information value of each node in the social network data; S3. Calculate the adaptive information entropy weight matrix for each node based on the node information entropy results; S4. Use graph neural networks to aggregate node features in a social network graph. During the aggregation process, the graph neural network uses an adaptive information entropy weight matrix to effectively aggregate the features of adjacent nodes to the target node and update the feature representation of the target node. S5. Based on the updated node feature representation, the graph neural network is trained using the variational information entropy optimization loss function, so that the information entropy of hot event propagation nodes tends to be maximized and the information entropy of ordinary nodes tends to be minimized, thereby obtaining the optimized graph neural network model. S6. The optimized graph neural network model is applied to the real-time social network data stream to construct a dynamic social network graph from the real-time social network data, and social network hot events are identified based on the node feature aggregation results and the adaptive information entropy weight matrix. S7. Output the results of social network hot topic identification, including the topic of the hot topic, key propagation nodes and propagation path of the event.
2. The method for identifying trending events on social networks based on graph neural networks according to claim 1, characterized in that, S1 includes the following steps: S11. Collect social network data, wherein the social network data includes user interaction data, content data, and timestamp information indicating the time when user behavior occurs, wherein the user interaction data includes follow relationships, forwarding behavior, and comment interaction between users, and wherein the content data includes text, image, and video information posted by users; S12. Analyze social network data and construct a social network graph based on user interaction data and timestamp information. The social network graph is represented as a directed weighted graph: G = (V, E, W); Where V is the set of nodes in the social network, and the nodes in the node set are... Let E be a user or information unit in the social network data, and let E be the set of edges in the social network. For users With users The interactions between them, including following, forwarding, and commenting, Let the set of edge weights in the social network graph be the weights in the set of edge weights. The weights, calculated based on user interaction frequency, interaction content similarity, or time decay factor, are used to measure the strength of the propagation relationship between nodes. ; in, Indicates user With users In social networks, the frequency of interaction matters; the higher the frequency, the greater the weight. Indicates user With users The similarity between the content is calculated using cosine similarity. The time decay function represents the time of interaction. Regarding the impact on edge weights, earlier interactions are assigned lower weights. , , This is a hyperparameter.
3. The method for identifying social network hotspot events based on graph neural networks according to claim 1, characterized in that, S2 includes the following steps: S21. Extract feature representations for each node in the social network graph, including user content features, user interaction features, and time-series features. The user content features are extracted based on text, image, or video information posted by users, and the user interaction features are based on edge weights. The extracted time series features are used to characterize the evolution pattern of user behavior over time. S22. Calculate the initial information entropy of each node in the social network graph, wherein the initial information entropy is used to measure the information distribution characteristics of the node in the social network: ; in, For nodes The initial information entropy, For nodes The set of neighboring nodes, For nodes Propagate information to neighboring nodes The probability of propagation; S23. Calculate the information diversity entropy of the node, wherein the information diversity entropy is used to characterize the category distribution of the information received by the node: ; in, For nodes The information diversity entropy, where C is the set of all information categories. For nodes The probability of receiving information of category c; S24. Calculate the information flow entropy of the nodes, which measures the uniformity of information propagation among nodes in the social network: ; in, For nodes Information flow entropy For nodes The total propagation weight; S25. Calculate the node information entropy results. Combine the node information entropy results with the initial information entropy, information diversity entropy, and information flow entropy to characterize the role of nodes in the propagation of hot events: ; in, For nodes The comprehensive information entropy, , , These are the weight parameters.
4. The method for identifying social network hotspot events based on graph neural networks according to claim 1, characterized in that, S3 includes the following steps: S31. Read the node information entropy result The node information entropy result is used to measure the information value of each node in the social network graph during the propagation of hot events. The larger the node information entropy result, the more significant the node's role in the propagation of hot events. S32. Calculate the adaptive information entropy normalization weights for each node in the social network graph. These normalization weights are used to map information entropy values at different scales to a unified interval, resulting in the normalized node information entropy. ; S33. Calculate the adaptive information entropy weights of the nodes. These weights are used to adjust the aggregation strength of neighbor features in the graph neural network. ; in, For nodes Adaptive information entropy weights, For adaptive adjustment coefficient, Use the Sigmoid activation function; S34. Calculate the feature aggregation weights of each node in the social network graph relative to its neighboring nodes. The feature aggregation weights are calculated by combining the adaptive information entropy weights and the edge weights of the neighboring nodes: ; in, For nodes For neighboring nodes Weights for feature aggregation; S35. Calculate the adaptive information entropy weight matrix that ultimately guides the aggregation of neighbor features in the graph neural network. : 。 5. The method for identifying social network hotspot events based on graph neural networks according to claim 4, characterized in that, S4 includes the following steps: S41. Read the feature representations of each node in the social network graph at the l-th layer of the graph neural network. ,in Represents a node In the feature vector of the l-th layer of the graph neural network, the feature vector captures the nodes. The basic attributes of trending events on social networks; S42. Using an adaptive information entropy weight matrix For nodes Aggregate the features of neighboring nodes: ; in, This is the information entropy modulation coefficient, used to amplify or attenuate the effect of normalized comprehensive information entropy on feature aggregation. As a parameter sensitive to feature similarity, Represents a node with neighboring nodes The squared Euclidean distance between the feature representations in the l-th layer. These are the residual connectivity coefficients, used to preserve nodes. In the feature information of the previous layer, This is the normalization constant.
6. The method for identifying social network hotspot events based on graph neural networks according to claim 1, characterized in that, S5 includes the following steps: S51. Based on the updated node feature representation The predicted information entropy of a node is calculated, and this entropy is used to measure the importance of a node in the propagation of hot events. ; in, For nodes Predictive information entropy, For nodes Propagate information to neighboring nodes Predicted propagation probability: ; in, For nodes and nodes Feature similarity at layer l+1 of the graph neural network: ; S52. Calculate variational information entropy to optimize the loss function, so that the information entropy of hot event propagation nodes tends to be maximized, and the information entropy of ordinary nodes tends to be minimized: ; in, To optimize the loss for variational information entropy, For nodes The true label is 1 for nodes that spread hot events, and 0 for ordinary nodes; S53. Gradient descent is used to optimize the graph neural network model. The parameters of the graph neural network model are updated by calculating the gradient of the loss function based on the variational information entropy, so that the graph neural network model is continuously optimized during training, thereby improving the performance of hotspot event recognition. ; in, Let be the parameters of the graph neural network model at the t-th iteration. For learning rate, The gradient of the loss function with respect to network parameters is optimized for variational information entropy.
7. The method for identifying social network hotspot events based on graph neural networks according to claim 1, characterized in that, S7 includes the following steps: S71. Predictive Information Entropy Based on Hotspot Event Propagation Nodes Identify the key dissemination nodes of trending events: ; Where K is the set of key propagation nodes, For nodes Predictive information entropy, For nodes with neighboring nodes Edge weights between them The information entropy threshold, These are the propagation intensity thresholds; only nodes that meet these two threshold conditions are identified as key nodes in the propagation of hotspot events. S73. Calculate the propagation path of hot events, wherein the propagation path is constructed based on the information flow between key propagation nodes: ; Where P represents the optimal propagation path of the trending event. Given the set of all possible propagation paths on the key propagation node set K, the optimal solution for each propagation path is based on the propagation weights of all edges along the path. and target node Predictive information entropy The weighted sum and maximization calculation is obtained; S74. Determine the event theme of the hot topic, which is calculated based on the content characteristics of the hot topic's propagation nodes: ; Where T represents the topic category of the trending event, and C represents the set of all topic categories. As a key propagation node The probability of belonging to topic category c is determined by maximizing the sum of the probabilities of propagation nodes in a certain category; S75. Output the identification results of social network hot events, the identification results including the event topic T, the set of key propagation nodes K, and the event propagation path P.