A method for predicting the popularity of network event tags based on multi-label influence

By building a deep learning model with multi-label impact, combined with graph attention network and graph capsule network, the accuracy of label popularity prediction in social networks is solved, and more efficient label popularity prediction is achieved, and prediction performance is improved.

CN115858899BActive Publication Date: 2025-07-11NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211605375.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-14
Publication Date
2025-07-11
Estimated Expiration
2042-12-14

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the popularity of network event tags in social networks, especially when multiple related tags are propagated simultaneously, and their joint dynamic propagation process cannot be effectively captured, and the existing methods rely on strong assumptions, resulting in insufficient prediction performance.

Method used

An end-to-end deep learning regression prediction model is adopted to construct event tag propagation relationship diagrams, global tag relationship diagrams and local impact attribute diagrams, combined with graph attention networks and graph capsule networks, learn semantic features and group indicators between labels, construct static and dynamic features, and predict the popularity of multi-label impact.

Benefits of technology

It improves the accuracy of label popularity prediction in social networks, can better reflect the mutual influence between labels, and significantly improves prediction performance, especially in the popularity prediction of hot events, which has increased by 25.9% and 29.3%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115858899B_ABST
    Figure CN115858899B_ABST
Patent Text Reader

Abstract

The present invention provides a method for predicting the popularity of network event tags based on multi-tag influence, which collects event tag propagation data and related user data related to the event; constructs a tag propagation relationship network, obtains node relationships and node attributes, and establishes a propagation popularity prediction model including: a feature aggregation component, including a static semantic feature aggregation and a dynamic group propagation feature aggregation process; a local aggregation component, composed of a graph capsule network, to learn the feature representation of local tag aggregation; a dynamic time series representation component, to learn the time series process of tag propagation evolution; the three components simulate the propagation influence process between tags, train the model; input the event tags to be predicted and the propagation influence network data related to the tags into the trained model, and output the popularity index that the concerned social network event tags may generate in the future. The present invention can predict the future popularity of the concerned social network event tags on social media.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of deep learning and social network public opinion analysis, and particularly relates to a method for predicting the popularity of network event tags based on multi-label influence. Background Art

[0002] As an extension of real-world events on the Internet, social media has become an important platform for the online dissemination of social event content. Official media or self-media publish real-time events occurring in society to social networks through social network services. After becoming network events through dissemination, they attract a wide range of users to participate and further have a social impact. Therefore, the popularity of network events is directly related to the impact of specific events in reality. Events in society are usually labeled with semantically general phrases (phrases starting with the # symbol). When entering the explosive dissemination situation, multiple tags will be used to divert and maintain it. Network events continue to spread under the mutual influence of these tags to achieve the effect of expanding the influence scope of social networks. In reality, if a network event has caused a certain impact on social media, for example, becoming a "hot search" event, multiple tags reflecting the semantic connotations of the event will often emerge as the event develops. The strong correlation between these tags related to the event promotes the scope of the event's dissemination on the network. The emergence of the social network event tag model has changed the mode of dissemination of social hot events on the network. Tags have stronger aggregation effects and semantic generalization capabilities, promoting network users to participate in the dissemination process. The dissemination of network hot events on social networks is usually accompanied by multiple tags, attracting traffic in the form of different tags as topics.

[0003] Predicting the popularity of social content is a hot research issue in network dissemination prediction. The research on popularity prediction can intuitively evaluate how much dissemination volume the events concerned by users will obtain in a future period of time, provide data support for understanding the dissemination law of social network events, and predict the dissemination scale of events through social media data. Since each sub-topic generated by a network event will be participated by user groups with different interests, and these topics represent the expressions of the groups towards the event, the dissemination popularity of the event tag set can be regarded as a measurement method for evaluating the group attention in the event. Analyzing the dissemination influence relationship among many tags under a network hot event and predicting the popularity by using the mutual influence of tags can provide support for specific tasks such as public opinion event monitoring, opinion and stance prediction.

[0004] A large amount of previous work has studied the prediction of the popularity of hashtags in social network events. On the one hand, these methods are usually based on macroscopic analysis and provide a cascade propagation model for individual hashtag items. However, they cannot directly capture the joint dynamic propagation process of a set of hashtag information items with the network event as a whole, so they cannot reflect the propagation impact generated by multiple relevant and complex hashtags during the outbreak period of the network event within the same time window. On the other hand, most of the existing analysis methods for the propagation impact of social network hashtags are based on point process time series models. This type of model usually only contains a single attribute of the event hashtag itself (such as the number of forwards, time series relationship, etc.), and uses a generative model to predict the future popularity of hashtags. This type of method usually relies on strong assumption conditions, has deficiencies in obtaining the potential relationships of hashtag features, and has certain limitations in prediction performance.

[0005] When the propagation of a network event is in a popular period, the generated hashtags can be regarded as a semantic representation with group aggregation, that is, different user group variables are hidden under different hashtags of the event. Therefore, these hashtags imply the propagation attributes of the group. Therefore, the present invention proposes an end-to-end deep learning regression prediction model, which performs deep learning modeling considering the mutual propagation influence between different hashtags in the event hashtag propagation stage, and introduces semantic features and two group indicators as key features affecting propagation, so as to more accurately predict the future popularity of target hashtags in the social network. Summary of the Invention

[0006] The present invention aims to provide a method for predicting the popularity of network event hashtags based on multi-hashtag influence to solve the existing problems.

[0007] The technical solution is as follows:

[0008] A method for predicting the popularity of network event hashtags based on multi-hashtag influence, comprising the following steps:

[0009] S1. During the propagation of real-world social network events, crawl the hashtags and texts related to the social network events;

[0010] Further, in S1, in the dataset of the propagation of real-world social network events, use equal topic words as keywords to crawl social texts for a certain period of time.

[0011] S2. Clean and organize the crawled data; and perform preprocessing on the relationship features preprocessing. Specifically, it includes:

[0012] S201. Data cleaning. The csv file after data cleaning records the publication status of tweets. The data fields include the user ID who published the tweet, the tweet content, the publication time, and the user IDs involved in tweet interactions (including retweets / original tweets / comments. For original tweets, the interaction ID is the same as the tweet ID). Each record represents the publication of a tweet related to an event;

[0013] S202. Preprocessing of relationship features. According to the characteristics of tweets in different datasets, the preprocessing process additionally extracts the general event title and establishes label relationships for tweets with multiple tags in one tweet.

[0014] S3. Collect event-related tags for network events and construct an event tag propagation relationship graph, a global label relationship graph G, and a local influence attribute graph G using the associations between tags i ; and calculate the propagation feature attributes of nodes. First, divide the time into windows based on observable data, and then construct a local propagation influence feature graph G of network event tag i based on the data in the time window i , including extracting a static semantic feature graph and a dynamic group propagation time series graph sequence

[0015] Specifically, it includes: S301. Construct an event tag propagation relationship graph. Two metrics are selected as the source of label relationships in the relationship graph. These two metrics are explicit relationships and implicit semantic relationships;

[0016] (A) Explicit relationship metric. If a user explicitly aggregates more than two tags in a tweet, it means that these tags have a propagation influence relationship during the propagation process. The specific formalization is as follows:

[0017]

[0018] where C O-occur (i,j) represents the frequency of two tags appearing in a tweet at the same time, and N represents the set of co-occurring nodes of all neighbors of node i. This formula expresses the association degree between tag i and tag j in the entire event corpus and can represent explicit propagation influence.

[0019] (B) Implicit semantic relationship metric. When there is a lack of explicit relationships among the network event tags after an outbreak, use the already extracted explicit tags and establish relationship associations through semantic similarity within the observation window. The purpose is to extract event tags that are semantically related but do not have an explicit # character marker. Specifically, the model uses the pointwise mutual information (PMI) method to establish links between the semantic relationships of tags. This method can express the weight relationship between tags in the event semantic data. The specific formalization is as follows:

[0020]

[0021]

[0022] Among them, d(i, j) is the total number of tweets in which observable window event labels i and j appear simultaneously. Here, different from explicit features, explicit features are explicit labels that are clearly marked with the # sign in the tweet, while d(i, j) here is the co-occurrence relationship that appears in the tweet text and does not contain #. d(i) and d(j) are the total number of tweets in the set that contain i and j at least once. D is the total number of tweets in the social network event. Generally speaking, a positive PMI value means that the labels in the event label library have a high semantic correlation.

[0023] (C) Establishing the relationship weights between social network event labels, and implicit semantic relationships are only established between label pairs with positive PMI values. Finally, the relationship weights between social network event labels are determined by the summation method:

[0024] R(i, j) = {R ex (i, j) + R im (i, j)}

[0025] For social network events, the above process can establish an event label relationship graph G = <V, E> in the data of the observable time window;

[0026] (D) Sampling a sub-network G i of a fixed size for the target label i from the global graph G of network event label relationships, rather than directly processing the global network G itself with a lot of noisy information. The sampling method for this step selects the restart random walk algorithm RWR. Using the RWR method, the relevant labels that are most likely to affect the propagation can be selected, and the upper limit of the associated label nodes is constrained, that is:

[0027] G i = RWR(G, i, m, R)

[0028] Where G represents the global relationship graph of event labels, m represents the constraint number of sampled nodes, i represents the target label, and R represents the weight R(i, j) between the nodes of the event label relationship graph.

[0029] S302. Construct a static semantic relationship attribute graph, and use semantic features as static attributes in the label relationship graph (label semantics is independent of time sequence) to construct an event label semantic relationship attribute graph G sem , and the process is as follows:

[0030] (A) Extraction of label semantic features. Since most labels are abbreviations of event summaries, especially in English datasets, it is difficult to directly obtain sufficient semantic information from the text of the labels themselves. The tweets where the labels are located are supplemented to obtain the semantic information of the labels within the observable time window. Therefore, the tweet s with the label that has the largest number of retweets within the observable time is obtained and used as the feature to explain the label semantics.

[0031] (B) Embedding representation of event label semantic features. After establishing the text library for semantic interpretation of the labels, the sentenceTransformer5 interface based on the bert model is called to perform semantic initial vector embedding on the labels, formalized as:

[0032] H s ={bert2sentence(i)}

[0033] where 0 ≤ i ≤ |V|, H s ∈R |v|×d , |V| represents the number of all label nodes, and d represents the embedding dimension;

[0034] (C) Construction of a static semantic relationship attribute graph. Using the event label propagation local influence relationship graph G in step 301 i and the node attributes obtained in the previous step to construct a static semantic relationship attribute graph where, V i is the set of associated nodes of the target label i, E i is the set of label node relationships, is the set of semantic feature representations of node V i , and

[0035] S303. Construct a sequence of event label propagation temporal attribute graphs. Based on the event label propagation temporal attribute graph This model performs segmentation based on temporal subgraphs, that is, retaining the valid nodes and edges within the time window t, and calculating the node group feature Ht of the influence propagation according to the group influence of the label nodes at different times, such that where 0 ≤ t < n, representing the serialized time window. Such a sampling method can reflect the dynamic influence process between different topic labels within the time window. This step is divided into four steps:

[0036] (A) Calculate the social influence of the participating group. Define the social influence of label i within the t window, and perform the following calculation for each label node:

[0037]

[0038] Where N t (n) represents the total number of labeled tweets in the time window t, N t (E) represents the total number of tweets of all associated label nodes in the subgraph under the time window t. Intuitively, this metric expresses the overall influence degree of the group participating in this label on the event label at time t.

[0039] (B) Calculate the influence of the participating group in the spread, and define the group spread influence of hashtagi at time t. The following calculations are performed for each node:

[0040]

[0041] Among them, represents the total number of followers of the users who posted tweets with the label at time t of the time window, represents the total number of users who posted event-related tweets under the time window t. This metric expresses the degree of group attention of the group participating in the event label to the current label.

[0042] (C) Construct the dynamic group attribute H inf (t), and select the dynamic group influence feature H inf (t) that is decisive for the spread as the node attribute. Since the spread influence of different labels is different within different time windows, two key metrics of label influence are used as the influence evaluation criteria for the label participating groups in the time series, and thus the dynamic spread attribute H inf (t) of the target label i is obtained, which is formalized as follows:

[0043] H inf (t) = [O(t), M(t)]

[0044] (D) Construct the time series graph sequence Construct the attribute graph G inf based on the dynamic attribute and dynamic relationship through the dynamic spread attribute H t , and then obtain the time series attribute graph sequence of the event label spread within each time window:

[0045]

[0046] Among them Then for the label i of each hot event in the dataset, the feature graph of the target label i is constructed according to the above method as the sample input of the deep learning model.

[0047] Step 4: Construct semantic features and temporal group features. The constructed semantic features and temporal group features include: a propagation relationship feature embedding layer based on a graph attention network, a feature fusion layer that fuses static and dynamic features, and a local feature aggregation layer based on a graph capsule network. The propagation relationship feature embedding layer based on the graph attention network further includes a static semantic feature aggregation representation and a dynamic propagation feature aggregation representation. For the static semantic relationship attribute graph, is input into the above-mentioned static semantic feature aggregation representation layer to obtain a static semantic node vector representation matrix H sem , and for the dynamic propagation temporal attribute graph sequence, for each propagation attribute graph in each time window is input into the dynamic propagation feature aggregation representation layer to obtain a dynamic propagation matrix H' composed of t windows within the observable time dym ; The static semantic node vector representation matrix H sem and the dynamic propagation matrix H' dym are input into the feature fusion layer to obtain a fusion feature H that fuses the semantic features and propagation features of the fusion label nodes f ; The propagation attribute graph is input into the local feature aggregation layer based on the graph capsule network to obtain a graph embedding vector h G .

[0048] Specifically, it includes:

[0049] S401. Learn the propagation relationship feature embedding layer. The propagation relationship feature embedding is mainly responsible for learning the propagation influence relationship between different labels, including static semantic feature aggregation representation and dynamic propagation feature aggregation representation. The graph neural network GAT, which can learn the influence relationship weights, is selected as the method for graph neural network feature representation. Different weight coefficients between nodes are learned during the supervised learning process of GAT, and the hidden mutual influence relationship is obtained in the node representation.

[0050] The input of GAT contains two parts, the node feature vector H ∈ R |V|×d and the adjacency matrix N t ∈ R |V|×|V| , d represents the dimension of the feature, and n represents the number of nodes in the subgraph. For each node vi and attribute hi in the graph, there is:

[0051] H = [h1, h2,..., h1,] T

[0052]

[0053]

[0054] H o= [h'1, h'2,..., h' n , T , H o ∈ R |V|×d'

[0055] where α ∈ R 2×d’ , Θ represents a trainable weight matrix, j ∈ N(i) indicates that there is an edge between label node j and i in the adjacency matrix (indicating that j and i are related), the LeakyReLu method is selected as the non - linear activation function, a i,j represents the weight of the mutual influence relationship between label node i and label node j, H O is the matrix composed of the output embedding vectors, |V| represents the number of graph nodes, d’ represents the dimension of the output node features, and || represents the concatenation operation.

[0056] For the static semantic relationship attribute graph, is input into the learning propagation relationship feature embedding layer to obtain the static semantic node vector representation matrix H sem , For the dynamic propagation time - series attribute graph sequence, for each propagation attribute graph in each time window is used as input to obtain the tensor H' dym composed of t windows within the observable time:

[0057]

[0058] Obviously where t represents the number of time windows, |V| represents the number of sub - graph nodes, and F'2 represents the dimension of the output node features.

[0059] S402. The feature fusion module is mainly responsible for fusing the semantic features and propagation features of the label nodes and then using them as the input for the next layer. To maintain the invariance of node semantics during the time series, the model broadcasts H sem in the matrix H' dym , that is, the feature representation layer is formalized as follows:

[0060] H f = H sem || H’ dym

[0061] || represents the concatenation operation, so

[0062] S403. According to the characteristic that there is a locally strong semantic correlation of tags as the event situation develops, these tags are in a state of strong intra-graph connection. In order to capture this strong local correlation between propagation graphs and better express the hierarchical structure relationship from local to global, inspired by the graph capsule network, the model proposed in the present invention applies a routing mechanism to vote on the group effect of tag nodes to better capture the local-to-global relationship in the graph, and then infers the local-to-global hierarchical relationship through multiple rounds of iteration, and finally obtains the temporal graph feature representation. Specifically, this process is mainly composed of a temporal attribute graph mapping to the graph embedding vector h G The process of this step is divided into three steps:

[0063] (A) Hierarchicalize the temporal graph and establish the relationship between the low-level local and high-level global of the temporal graph through a voting mechanism. It is stipulated that v represents the capsule node in the low-level voting graph, and u represents the capsule node in the high-level routing graph. First, use the feature fusion layer and the temporal graph as the initial voting matrix, that is:

[0064]

[0065] where N t represents the adjacency matrix of the low-level voting graph, represents the voting representation vector of the initial bottom layer network in the graph capsule network, and each vector is regarded as the voting weight representation of v i for the capsule node u j in the high-level cluster, |N| represents the number of nodes in the capsule, and F'3 = F'1 + F'1 represents the dimension of the input vector.

[0066] (B) Establish dynamic routing selection. The task of this process is to iteratively calculate the routing weight C i,j from the low-level node v to the higher-level capsule u, that is, which group of low-level v j can activate the high-level cluster node u j (u j can be regarded as a more closely related relationship cluster in the subgraph) to obtain the local-to-global activation relationship. Then, the vote in step (A) is weighted and calculated to obtain the routing weight C j from b i,j in the low-level graph to the locally aggregated high-level graph, that is:

[0067]

[0068] is initialized to 0. Then, the dynamic routing weight is calculated through R iterations using the following three formulas The three formulas are as follows:

[0069] bi,j = b i,j + v j|i .u j

[0070]

[0071]

[0072] Among them, the role of squash (non-linear "squeeze" function) is to calculate the routing possibility of node v in capsule u j , that is, the probability that node v votes for u j|i . Among them, the role of v i|j ·u j is to calculate the consistency between each group of votes and the high-level capsule, so that it is possible to focus more on aggregating information from neighbors who may be in the same cluster. After R iterations, this process obtains the capsule nodes u of the high-level aggregated graph and the adjacency matrix of the high-level abstraction, which is expressed as: i|j ·u j

[0073] G route = (A, u)

[0074] A = C T N C, A ∈ R |V|×|U|

[0075] Among them, N is the adjacency matrix in the low-level voting graph, |V| represents the number of nodes in the low-level voting graph, |U| represents the number of nodes in the high-level routing graph, and C represents the routing weight matrix from the low level to the high level composed of C i,j , C ∈ R |V|×|U| , so A can be regarded as the adjacency matrix of the high-level routing graph, and u is the feature vector of the nodes in the high-level routing graph, which is obtained from the previous formula. Therefore, the above process can be simplified to the following transformation:

[0076] U, A = Route(Vote(V, N))

[0077] That is:

[0078] U, A = RV(V, N)

[0079] (C) Establish a temporal attribute graph representation, and then repeat the above steps A and B to abstract the high-level cluster graph to the representation of the whole graph embedding again. The role of doing this is to maximize the retention of the characteristics of the influence of local hashtag tags on propagation on the basis of the graph capsule network, that is:

[0080]

[0081] ​The "1" in the formula represents the graph representation that only aggregates one node after being abstracted to a higher level, and the feature vector of this node is represented as the graph representation propagation influence representation vector under the current time window t. Then, for each temporal attribute graph, the local aggregation layer is used for the graph representation process:

[0082]

[0083] S5. Input the representation vectors h of different temporal subgraphs in the sample t into the LSTM model for dynamic temporal representation learning, and then input the result into the fully connected layer to obtain the prediction result. Use the error between the prediction result and the obtained true value label of the sample to guide the model learning, which is divided into three steps:

[0084] S501. Then, through the representation vectors h of different temporal subgraphs in the sample t , in order to utilize these temporal features and better obtain the propagation influence brought by the feature changes in the time series, a long short-term memory (LSTM) kernel is applied in this part. The specific calculation formula is as follows:

[0085]

[0086]

[0087]

[0088]

[0089]

[0090] h t = tanh(c i ) * o i

[0091] where h t is the implicit feature output at the t-th moment, represents the Hadamard product, U j , W j , b j , j ∈ ({z, f, o, c}) are learnable parameters, z i , f i and o i are the forget gate vector, input gate vector, and output gate vector of the feature in the t-th window respectively. Finally, the result at the t + 1 moment is predicted through the fully connected layer:

[0092] Δy’ = σ(Wh t )

[0093] S502. Obtain the true value label of the sample. For the local propagation influence feature map G of the network event label i i , count the of this sample within the t+1 snapshot i . Due to the particularity of label propagation in the social network, original tweets with labels are also regarded as behaviors of propagation. Therefore, the present invention uses the sum of the number of retweets and the number of original tweets in the form with labels as the attention index of the user group, that is where represents how many people have retweeted the tweet with the event label, represents how many original tweets cover the target label.

[0094] S503. For the prediction of network popularity, the popularity prediction is regarded as a regression model. Therefore, MLSE is used in the text as the target loss function based on the regression model:

[0095]

[0096] where Δy' represents the predicted popularity index, and y i represents the actual propagation index.

[0097] A computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the method for predicting the popularity of network event labels based on multi-label influence as described above is implemented.

[0098] A computer-readable storage medium stores a program. When the program is executed by a processor, the method for predicting the popularity of network event labels based on multi-label influence as described above is implemented.

[0099] The present invention establishes a method for predicting the popularity of network event tags based on multi-tag influence. Specifically, first, considering the mutual propagation influence among tags related to an event during the network propagation process, an event tag propagation relationship graph, a global tag relationship graph, and a local influence attribute graph are constructed by collecting event-related tags and their associations. Then, since the propagation process of social network events is dynamically changing and the generated event tags exhibit a semantic aggregation process and an evolution process, a semantic feature and a temporal group feature are constructed using the propagation relationship feature embedding layer of the graph attention network, the feature fusion layer that combines static and dynamic features, and the local feature aggregation layer based on the graph capsule network, providing more reliable accuracy for predicting the popularity of network event tags. Finally, dynamic temporal representation learning is performed through a temporal model to predict the popularity of tags for hot events. For different social platforms, a prediction model for the popularity of network event tags based on multi-tag influence can be trained according to the data, better solving the problem of predicting the popularity of hot events. It can be used to predict events with a large impact on public opinion, such as social hot issues, public opinion events, stance events, opinion events, diplomatic events, and international events, etc.

[0100] Compared with the prior art, the present invention has the following technical effects:

[0101] 1. The present invention designs a feature aggregation component based on the graph attention network for obtaining the hidden mutual influence relationship in the node representation during the propagation process of social network blog posts. It includes a static semantic feature aggregation and a dynamic group propagation feature aggregation process, introducing semantic features and two group metrics as key features affecting propagation. This component models the association between tags and the internal semantic relationship of tags.

[0102] 2. The present invention provides a method using the graph capsule network to perform representation learning on the group aggregation characteristics behind the combination of static semantic features and dynamic temporal features in view of the locally strong correlation of tag semantics in the development of event situations, capturing the strong local correlation between such propagation graphs, better expressing the hierarchical structure relationship from local to global, and thus modeling the propagation influence network to reflect the mutual influence relationship of groups under different tags on event propagation.

[0103] 3. Since the propagation process of social network events is dynamically changing and the generated event tags exhibit a semantic aggregation process and an evolution process, the present invention uses the LSTM temporal model to learn the feature representation of the propagation evolution process, calculates the semantic correlation and the structural correlation, simulates the semantic aggregation process and the evolution process of event tags, learns the potential features of the propagation influence between different tags under the temporal process, and further predicts the future popularity of target tags.

[0104] 4 The present invention addresses the problem of the continuous increase in the nodes of the label relationship network diagram, which may generate a large number of noisy labels. It uses the random walk algorithm to downsample the label subgraph with strong influence, screens the key label set of influence, constructs a time-series-based event label relationship network subgraph, avoids node noise in the global graph, and reduces the computational complexity of deep learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0105] Figure 1 It is a schematic diagram of the steps of the method for predicting the popularity of network event labels affected by multiple labels of the present invention;

[0106] Figure 2 It is a flowchart of the steps of the method for predicting the popularity of network event labels affected by multiple labels of the present invention;

[0107] Figure 3 It is the internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0108] The following will describe in detail the embodiments of the present invention in conjunction with the drawings and embodiments, so as to fully understand how the present invention uses technical means to solve technical problems and achieve the implementation process of technical effects and implement accordingly. It should be noted that as long as there is no conflict, the various embodiments in the present invention and the various features in each embodiment can be combined with each other, and the formed technical solutions are all within the protection scope of the present invention.

[0109] In addition, the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0110] See Figure 1 and Figure 2 , a method for predicting the popularity of network event labels based on multi-label influence of the present invention includes at least the following steps:

[0111] Step 1: In a large real-world social network event propagation dataset, crawl the labels and texts related to social network events;

[0112] Step 2: Clean and organize the crawled data; and preprocess the relationship features.

[0113] Step 3: Collect event-related labels for network events, and use the associations between the labels to construct an event label propagation relationship graph, a global label relationship graph G, and a local influence attribute graph G i; and the propagation characteristic attributes of the computing nodes. First, divide the time based on the observable data into windows, and then construct the local propagation influence feature map G of the network event label i according to the data under the time window i , including extracting the static semantic feature map and the dynamic group propagation time series graph sequence

[0114] Step 4: Construct semantic features and temporal group features. The constructed semantic features and temporal group features include: a propagation relationship feature embedding layer based on a graph attention network, a feature fusion layer that fuses static and dynamic features, and a local feature aggregation layer based on a graph capsule network. Among them, the propagation relationship feature embedding layer based on the graph attention network further includes a static semantic feature aggregation representation and a dynamic propagation feature aggregation representation. For the static semantic relationship attribute graph, is input into the above-mentioned static semantic feature aggregation representation layer to obtain the static semantic node vector representation matrix H sem . For the dynamic propagation time series attribute graph sequence, for each propagation attribute graph under each time window is input into the dynamic propagation feature aggregation representation layer to obtain the dynamic propagation matrix H' composed of t windows within the observable time dym ; The static semantic node vector representation matrix H sem and the dynamic propagation matrix H' dym are input into the feature fusion layer to obtain the fusion feature H that fuses the semantic features and propagation features of the fusion label nodes f ; The propagation attribute graph is used as the input to the local feature aggregation layer based on the graph capsule network to obtain the graph embedding vector h G .

[0115] Step 5: Input the representation vectors h t of different temporal subgraphs in the sample into the LSTM model for dynamic temporal representation learning, and then input the result into the fully connected layer to obtain the prediction result. Use the error between the prediction result and the obtained sample true value label to guide the model learning

[0116] Specifically, in an embodiment of the present invention, it includes the following steps

[0117] Step 1: In the propagation of social network events in the real world, use the topic words as keywords to crawl social texts for a certain period of time

[0118] Step 2: Clean and sort the crawled data; and perform preprocessing on the relationship feature preprocessing. Specifically, it includes

[0119] Step 201: Data cleaning. The csv file after data cleaning records the publication status of tweets. The data fields include the user id of the tweet publisher, the tweet content, the publication time, and the user id of the tweet interaction (including retweet / original / post comment, and the interaction id of the original is the same as the publication id). Each record represents the publication of a tweet related to an event;

[0120] Step 202: Preprocessing of relationship features. According to the characteristics of tweets in different datasets, the preprocessing process additionally extracts the general event title and establishes label relationships for tweets with multiple tags in one tweet;

[0121] Step 3: Collect event-related tags for network events and construct an event label propagation relationship graph, a global label relationship graph G, and a local influence attribute graph G using the associations between the tags i ; and calculate the propagation feature attributes of the nodes. First, divide the time based on observable data into time windows, and then construct the local propagation influence feature graph G of the network event label i according to the data under the time window i , including extracting the static semantic feature graph and the dynamic group propagation time series graph sequence

[0122] Step 301: Construct an event label propagation relationship graph. Two metrics are selected as the source of label relationships in the relationship graph, and these two metrics are explicit relationships and implicit semantic relationships;

[0123] Step A: Explicit relationship metric. If a user explicitly aggregates more than two tags in one tweet, it means that these tags have a propagation influence relationship during the propagation process. The specific formalization is as follows:

[0124]

[0125] where C o-occur (i,j) represents the frequency of two tags appearing in one tweet at the same time, and N represents the set of co-occurring nodes of all neighbors of node i. This formula expresses the association degree between tag i and tag j in the entire event corpus and can represent the explicit propagation influence.

[0126] Step B: Implicit semantic relationship metric. When there is a lack of explicit relationships for the network event tags after the outbreak, use the already extracted explicit tags and establish a relationship association through the semantic similarity within the observation window. The purpose is to extract event tags that are semantically related but do not have an explicit # character marker. Specifically, the model uses the pointwise mutual information (PMI) method to establish a link for the semantic relationship between tags, and this method can express the weight relationship between tags in the event semantic data. The specific formalization is as follows:

[0127]

[0128]

[0129] Among them, d(i, j) is the total number of tweets in which observable window event labels i and j appear simultaneously. Here, different from explicit features, explicit features are explicit labels that are clearly marked with the # sign in the tweet, while d(i, j) here is the co-occurrence relationship that appears in the tweet text and does not contain #. d(i) and d(j) are the total number of tweets in the set that contain i and j at least once. D is the total number of tweets in the social network event. Generally speaking, a positive PMI value means that the labels in the event label library have a high semantic correlation.

[0130] Step C: Establish the relationship weight between social network event labels. Implicit semantic relationships are only established between label pairs with positive PMI values. Finally, the relationship weight between social network event labels is determined by the summation method:

[0131] R(i, j) = {R ex (i, j) + R im (i, j)}

[0132] For social network events, the above process can establish an event label relationship graph G = <V, E> in the data of the observable time window;

[0133] Step D: Sample a sub-network G i with a fixed size from the global graph G of network event label relationships, rather than directly processing the global network G itself with a lot of noisy information. The sampling method in this step selects the restart random walk algorithm RWR. Using the RWR method, the relevant labels that are most likely to affect the spread can be selected, and the upper limit of the associated label nodes is constrained, that is:

[0134] G i = RWR(G, i, m, R)

[0135] Where G represents the global relationship graph of event labels, m represents the number of constraints on the sampled nodes, i represents the target label, and R represents the weight R(i, j) between the nodes of the event label relationship graph. In this embodiment, m = 15;

[0136] Step 302: Construct a static semantic relationship attribute graph, and construct an event label semantic relationship attribute graph G sem by taking semantic features as static attributes (label semantics are independent of time sequence) in the label relationship graph. The process is as follows:

[0137] Step A: Extraction of label semantic features. Since most labels are abbreviations of event summaries, especially in English datasets, it is difficult to directly obtain sufficient semantic information from the text of the labels themselves. The tweets where the labels are located are supplemented with the semantic information of the labels within the observable time window. Therefore, the tweet s with the most retweets containing the label within the observable time are obtained and used as the features to explain the label semantics.

[0138] Step B: Embedding representation of event label semantic features. After establishing the text library for semantic interpretation of the labels, the sentence Transformer5 interface based on the bert model is called to perform semantic initial vector embedding on the labels, formalized as:

[0139] H s ={bert2sentence(i)}

[0140] where 0 ≤ i ≤ |V|, H s ∈R |v|×d , |V| represents the number of all label nodes, d represents the embedding dimension. In this embodiment, d = 32 and |V| = 20;

[0141] Step C: Construction of a static semantic relationship attribute graph. Using the event label propagation local influence relationship graph G in step 301 i and the node attributes obtained in the previous step to construct a static semantic relationship attribute graph where, V i is the set of associated nodes of the target label i, E i is the set of label node relationships, is the set of semantic feature representations of node V i , and

[0142] Step 303: Construction of a sequence of event label propagation temporal attribute graphs. Based on the event label propagation temporal attribute graph This model performs segmentation on the temporal subgraph, that is, retaining the valid nodes and edges within the time window t, and calculating the node group feature Ht of the influence propagation according to the group influence of the label nodes at different times, such that where 0 ≤ t < n represents the serialized time window. Such a sampling method can reflect the dynamic influence process between different topic labels within the time window. This step is divided into two steps:

[0143] Step A: Calculate the social influence of the participating group. Define the social influence of label i within the t window, and perform the following calculation for each label node:

[0144]

[0145] where N t (n) represents the total number of labeled tweets in the time window t, and N t (E) represents the total number of tweets of all associated label nodes in the subgraph under the time window t. Intuitively, this metric expresses the overall influence degree of the group participating in this label on the event label at time t.

[0146] Step B: Calculate the influence of the participating group in the spread, and define the influence of the group spread of hashtagi at time t. The following calculations are performed for each node:

[0147]

[0148] where represents the sum of the number of followers of the users who post tweets with the label at time t of the time window, represents the total number of users who post event-related tweets under the time window t. This metric expresses the degree of group attention of the group participating in the event label to the current label.

[0149] Step C: Construct the dynamic group attribute H inf (t), and select the key dynamic group influence feature H inf (t) as the node attribute. Since the spread influence of different labels is different in different time windows, two key metrics of label influence are used as the evaluation criteria for the influence of the group participating in the label under time series, and thus the dynamic spread attribute H inf (t) of the target label i is obtained, which is formalized as follows:

[0150] H inf (t) = [O(t), M(t)]

[0151] Step D: Construct the time series graph sequence Construct the attribute graph G inf based on the dynamic attribute and dynamic relationship through the dynamic spread attribute H t , and then obtain the time series attribute graph sequence of the event label spread in each time window:

[0152]

[0153] where Then, for the label i of each hot event in the dataset, the feature graph of the target label i is constructed according to the above method as the sample input of the deep learning model.

[0154] Step 4: Construct semantic features and temporal group features. The constructed semantic features and temporal group features include: a propagation relationship feature embedding layer based on a graph attention network, a feature fusion layer that fuses static and dynamic features, and a local feature aggregation layer based on a graph capsule network. Among them, the propagation relationship feature embedding layer based on the graph attention network further includes a static semantic feature aggregation representation and a dynamic propagation feature aggregation representation. For the static semantic relationship attribute graph, is used as the input to the above-mentioned static semantic feature aggregation representation layer to obtain a static semantic node vector representation matrix H sem . For the dynamic propagation temporal attribute graph sequence, for each propagation attribute graph in each time window is used as the input to the dynamic propagation feature aggregation representation layer to obtain a dynamic propagation matrix H' composed of t windows within the observable time dym ; The static semantic node vector representation matrix H sem and the dynamic propagation matrix H' dym are used as the input to the feature fusion layer to obtain a fusion feature H that fuses the semantic features and propagation features of the fusion label nodes f ; The propagation attribute graph is used as the input to the local feature aggregation layer based on the graph capsule network to obtain a graph embedding vector h G .

[0155] Specifically, it includes:

[0156] Step 401, Learn the propagation relationship feature embedding layer. The propagation relationship feature embedding is mainly responsible for learning the propagation influence relationship between different labels, including static semantic feature aggregation representation and dynamic propagation feature aggregation representation. The graph neural network GAT that can learn the influence relationship weights is selected as the method for graph neural network feature representation. In the supervised learning process of GAT, different weight coefficients between nodes are learned, and the hidden mutual influence relationship is obtained in the node representation.

[0157] The input of GAT contains two parts, the node feature vector H ∈ R V|×d and the adjacency matrix N t ∈ R |V|×|V| , d represents the dimension of the feature, and n represents the number of nodes in the subgraph. For each node v i and the attribute h i in the graph, there is:

[0158] H = [h1, h2,..., h1,] T

[0159]

[0160]

[0161] H O = [h'1, h'2,..., h' n , ] T , H o ∈R |V|×d'

[0162] where α ∈ R 2×d’ , Θ represents a trainable weight matrix, j ∈ N(i) indicates that there is an edge between label node j and i in the adjacency matrix (indicating that j and i are related), the LeakyReLu method is selected as the non-linear activation function, a i,j represents the weight of the mutual influence relationship between label node i and label node j, H O is a matrix composed of the output embedding vectors, |V| represents the number of graph nodes, d' represents the dimension of the output node features, and || represents the concatenation operation. In this implementation case, |V| = 20, d' = 64;

[0163] For the static semantic relationship attribute graph, is input into the learning propagation relationship feature embedding layer to obtain the static semantic node vector representation matrix H sem , For the dynamic propagation time series attribute graph sequence, for each propagation attribute graph under each time window is used as the input to obtain the tensor H' composed of t windows within the observable time: dym :

[0164]

[0165] Obviously where t represents the number of time windows, |V| represents the number of subgraph nodes, and F'2 represents the dimension of the output node features. In this implementation case, |V| = 20, t = 8, F'1 = 32, F'2 = 32;

[0166] Step 402: The feature fusion module is mainly responsible for fusing the semantic features and propagation features of the label nodes and then using them as the input for the next layer. To maintain the invariance of node semantics during the time series, the model broadcasts H sem in the matrix H' dym , that is, the feature representation layer is formalized as follows:

[0167] H f = H sem || H' dym

[0168] || represents the concatenation operation, so there is In this implementation case, |V| = 20, t = 8, F'1 = 32, F’2 = 32;

[0169] Step 403: According to the characteristic that the semantic of tags shows local strong correlation in the development of the event situation, these tags are in a state of strong connection within the graph. In order to capture the strong local correlation between such propagation graphs and better express the hierarchical structure relationship from local to global, inspired by the graph capsule network, the model proposed in the present invention applies a routing mechanism to vote on the group effect of tag nodes, so as to better capture the local-to-global relationship in the graph. Then, through multiple rounds of iteration, the hierarchical relationship from local to global is inferred, and finally the temporal graph feature representation is obtained. Specifically, this process is mainly from the temporal attribute graph mapped to the graph embedding vector h G The process of this step is divided into three steps:

[0170] Step A: Hierarchicalize the temporal graph and establish the relationship between the local part of the lower layer and the global part of the higher layer of the temporal graph through a voting mechanism. It is stipulated that v represents the capsule node in the lower-layer voting graph, and u represents the capsule node in the higher-layer routing graph. First, use the feature fusion layer and the temporal graph as the initial voting matrix, that is:

[0171]

[0172] where N t represents the adjacency matrix of the lower-layer voting graph, represents the voting representation vector of the initial bottom layer network in the graph capsule network, and each vector is regarded as the voting weight representation of v i for the capsule node u j in the higher-layer cluster, |N| represents the number of nodes in the capsule, and F'3 = F'1 + F'1 represents the dimension of the input vector. In this implementation case, |N| = 10, F'1 = 32, F’2 = 32; F’3 = 64;

[0173] Step B: Establish dynamic routing selection. The task of this process is to iteratively calculate the routing weight C i,j from the lower-layer node v to the higher-layer capsule u, that is, which group of lower-layer v j can activate the higher-layer cluster node u j (u j can be regarded as a more closely related relationship cluster in the subgraph) to obtain the activation relationship from local to global. Then, the voting in Step A is weighted and calculated to obtain the routing weight C j from v i,j in the lower-layer graph to the locally aggregated higher-layer graph, that is:

[0174]

[0175] is initialized to 0. Then, the dynamic routing weight is calculated through R iterations using the following three formulas The three formulas are as follows:

[0176] b i,j = b i,j + v j|i .u j

[0177]

[0178]

[0179] Among them, the role of squash (nonlinear "squeezing" function) is to calculate the routing possibility of node v in capsule u j That is, the probability that node v votes for u, where v j|i ·u i|j is used to calculate the consistency between each group of votes and the high-level capsule, so that it is possible to focus more on aggregating information from neighbors who may be in the same cluster. After R iterations, this process obtains the capsule nodes u of the high-level aggregated graph and the adjacency matrix of the high-level abstraction, expressed as: j i|j j ·u j G

[0180] G route = (A, u)

[0181] A = C T N, C, A ∈ R |V|×|U|

[0182] Among them, N is the adjacency matrix in the low-level voting graph, |V| represents the number of nodes in the low-level voting graph, |U| represents the number of nodes in the high-level routing graph, C represents the routing weight matrix from the low level to the high level composed of C i,j C ∈ R |V|×|U| Therefore, A can be regarded as the adjacency matrix of the high-level routing graph, and u is the feature vector of the nodes in the high-level routing graph, obtained from the previous formula. Therefore, the above process can be simplified to the following transformation:

[0183] U, A = Route(Vote(V, N))

[0184] That is:

[0185] U, A = RV(V, N)

[0186] Step C: Establish a temporal attribute graph representation, and then repeat the above Steps A and B to abstract the high-level cluster graph to the representation of the whole graph embedding again. The role of this is to maximize the retention of the characteristics of the influence of local hashtag tags on propagation on the basis of the graph capsule network, that is:

[0187]

[0188] In the formula, 1 represents the graph representation that only aggregates one node after being abstracted to a higher layer, and the feature vector of this node is represented as the graph representation influence vector of the label relationship under the current time window t. Then, for each time-series attribute graph, the local aggregation layer is used to perform the graph representation process:

[0189]

[0190] Step 5: Input the representation vectors h of different time-series subgraphs in the sample t into the LSTM model for dynamic time-series representation learning, and then input the result into the fully connected layer to obtain the prediction result. Use the error between the prediction result and the obtained true value label of the sample to guide the model learning, which is divided into three steps:

[0191] S501. Then, through the representation vectors h of different time-series subgraphs in the sample t , in order to utilize these time-series features and better obtain the propagation influence brought by the feature changes in the time series, the long short-term memory (LSTM) kernel is applied in this part. The specific calculation formula is as follows:

[0192]

[0193]

[0194]

[0195]

[0196]

[0197] h t =tanh(c i )*o i

[0198] where h t is the implicit feature output at the t-th moment, represents the Hadamard product, U j , W j , b j , j∈({z,f,o,c}) are learnable parameters, z i , f i and o i are the forget gate vector, input gate vector and output gate vector of the feature of the t-th window respectively. Finally, the result at the t + 1 moment is predicted through the fully connected layer:

[0199] Δy’ = σ(Wh t )

[0200] S502. Obtain the true value label of the sample, and for the local propagation influence feature map G of the network event label i i , count the of this sample within the t+1 snapshot and set this value as the true value label y of the sample i . Due to the particularity of label propagation in the social network, original tweets with labels are also regarded as behaviors of propagation. Therefore, the present invention uses the sum of the number of retweets and the number of original tweets in the form with labels as the attention index of the user group, that is where represents how many people have retweeted the tweet with the event label, represents how many original tweets cover the target label.

[0201] S503. For the prediction of network popularity, regard the popularity prediction as a regression model. Therefore, MLSE is used as the target loss function based on the regression model in the text:

[0202]

[0203] where Δy' represents the predicted popularity index, and y i represents the actual propagation index.

[0204] Such an architecture has two advantages:

[0205] (1) Better feature modeling ability. Considering the propagation influence mechanism between event labels, studying the data law of tweet propagation, designing a construction method for the propagation association relationship of event labels, then extracting the group index behind the event label propagation and the semantic features that trigger label aggregation, and designing a network event label popularity prediction model based on multi-label influence for the labels that become hot events in the social network. This model considers the mutual influence relationship of labels in network propagation and predicts the propagation popularity in the social network through the semantics and group relationships hidden in the labels.

[0206] (2) More accurate label popularity prediction ability. There is a significant performance improvement in the task of event label popularity prediction. At the same time, experiments verify that the core index MLSE of event labels in the dataset during the propagation process exceeds the existing optimal benchmark popularity prediction models. Compared with the optimal baseline model, it has increased by 25.9% and 29.3% respectively on the two instantiated datasets, which indicates that the proposed label propagation influence relationship and semantic features are very helpful for the popularity prediction model, and proves that the proposed model is superior in performance, and the mutual propagation influence has a significant impact on popularity, indicating that the assumptions proposed by the model are reliable and effective.

[0207] This embodiment utilizes the characteristics of tags, static semantics, dynamic groups, and semantic aggregation in information dissemination, and has a more accurate prediction for network event tags affected by multiple tags. Therefore, for different social texts, the parameters of the deep learning model with different pertinences can be adjusted to better solve the prediction problems within the semantic scope, such as the prediction of the spread of viewpoints and positions and the prediction of public opinion events.

[0208] The method provided in this embodiment can be used for online public opinion event prediction, prediction of the spread of viewpoints and positions, public opinion event monitoring, rumor monitoring, and emergency prevention of public events. In particular, it can be used for the prediction of hot events affected by multiple tags in social networks, such as public opinion hot events, viewpoint hot events, position hot events, etc. It can also be used for the network information supervision of enterprises to predict whether the information released by enterprises will be widely spread in the future.

[0209] In an embodiment of the present invention, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method for predicting the popularity of network event tags based on multiple tags as described above is implemented.

[0210] This computer device can be a terminal, and its internal structure diagram can be as Figure 3 shown. This computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a bus. Among them, the processor of this computer device is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of this computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the method for predicting the popularity of network event tags based on multiple tags is implemented. The display screen of this computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of this computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the shell of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0211] The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. Among them, the memory is used to store programs, and the processor executes the programs after receiving execution instructions.

[0212] The processor can be an integrated circuit chip with the ability to process signals. The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc. The processor can also be other general-purpose processors, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0213] Those skilled in the art can understand that Figure 3 the structure shown in [the figure] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0214] In an embodiment of the present invention, there is also provided a computer-readable storage medium, on which a program is stored, and characterized in that: when the program is executed by a processor, it implements the method for predicting the influence of a social network based on a heterogeneous network as described above.

[0215] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a computer device, or a computer program product. Therefore, the embodiments of the present invention can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0216] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, computer devices, or computer program products according to the embodiments of the present invention. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing terminal devices generate a device for realizing the functions specified in the flowcharts and / or block diagrams.

[0217] These computer program instructions can also be stored in a computer-readable memory capable of guiding the computer or other programmable data processing terminal devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in the flowchart.

[0218] The above has introduced in detail the application of the method for predicting the popularity of network event tags based on multi-tag influence, computer device, and computer-readable storage medium provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for predicting the popularity of network event tags based on multi-label influence, characterized in that It includes the following steps: S1. Propagate in the real-world social network events, and crawl the tags and texts related to the social network events; S2. Clean and sort out the crawled data; And preprocess the relationship features; Specifically including: S201. Data cleaning. The csv file after data cleaning records the publication situation of tweets. The data fields include the user id who publishes the tweet, the tweet content, the publication time, and the user id of the tweet interaction; each record represents the publication of a tweet related to the event; S202. Relationship feature preprocessing. According to the characteristics of tweets in different datasets, the preprocessing process additionally extracts the general event title and establishes label relationships for tweets with multiple tags in one tweet; S3. Collect event-related tags for network events, and construct an event tag propagation relationship graph, a global tag relationship graph G, and a local influence attribute graph G using the associations between the tags i ; and calculate the propagation characteristic attributes of the nodes. First, divide the time based on the observable data into time windows, and then construct a local propagation influence characteristic graph G of the network event tag i according to the data under the time window i , including extracting a static semantic feature graph and a dynamic group propagation time series graph sequence S4. Construct semantic features and temporal group features; The constructed semantic features and temporal group features include: a propagation relationship feature embedding layer based on a graph attention network, a feature fusion layer that fuses static and dynamic features, and a local feature aggregation layer based on a graph capsule network. Among them, the propagation relationship feature embedding layer based on a graph attention network includes a static semantic feature aggregation representation and a dynamic propagation feature aggregation representation; For the static semantic relation attribute graph, is used as the input to the above-mentioned static semantic feature aggregation representation layer to obtain the static semantic node vector representation matrix H sem . For the dynamic propagation time series attribute graph sequence, for each propagation attribute graph in each time window is used as the input to the dynamic propagation feature aggregation representation layer to obtain the dynamic propagation matrix H' composed of t windows within the observable time dym ; Take the static semantic node vector representation matrix H sem and the dynamic propagation matrix H' dym as inputs to the feature fusion layer to obtain the fused feature H that fuses the semantic feature and the propagation feature of the fused label node f ; Take the propagation attribute graph as an input to the local feature aggregation layer of the graph capsule network to obtain the graph embedding vector h G ; S5. Input the representation vectors h of different time-series subgraphs in the sample t into the LSTM model for dynamic time-series representation learning, then input the result into the fully connected layer to obtain the prediction result, and use the error between the prediction result and the obtained true value label of the sample to guide the model learning.

2. The method for predicting the popularity of network event tags based on multi-tag influence according to claim 1, wherein S3 specifically includes: S301. Construct an event label propagation relationship graph, and select two indicators as the source of label relationships in the relationship graph. These two indicators are explicit relationships and implicit semantic relationships; S302. Construct a static semantic relationship attribute graph, and use semantic features as static attributes in the label relationship graph to construct an event label semantic relationship attribute graph G sem ; S303. Construct a sequence of event label propagation temporal attribute graphs, based on the event label propagation temporal attribute graph This model is based on temporal subgraph segmentation, that is, nodes that are valid within the time window t are retained and edges And according to the group influence of label nodes at different times, calculate the group characteristics Ht of nodes where the influence propagates, such that where 0 ≤ t < n, representing the serialized time window.

3. A method for predicting the popularity of network event tags based on multi-tag influence according to claim 2, characterized in that S301 specifically includes: (A) Explicit relationship indicator. If a user explicitly aggregates more than two tags in a tweet, it means that these tags have a propagation influence relationship during the propagation process; the specific formalization is as follows: Among them, C o-occur (i, j) represents the frequency of the simultaneous occurrence of two tags in a tweet, and N represents the co-occurrence node set of all neighbors of node i. This formula expresses the degree of association between tag i and tag j in the entire event corpus, representing the explicit propagation influence; (B) Implicit semantic relationship indicator. When there is a lack of explicit relationships among the labels of the network events after the outbreak, use the already extracted explicit labels and establish relationship associations through the semantic similarity within the observation window, and extract event labels that are semantically related but do not have an explicit # character mark; Specifically, the model uses the pointwise mutual information PMI method to establish links for the semantic relationships between tags, and the specific formalization is as follows: Among them, d(i,j) is the total number of tweets in which event tags i and j appear simultaneously in the observable window. Here, different from the explicit features, the explicit features are explicit tags that are clearly marked with # in the tweet, while d(i,j) here is the co-occurrence relationship that appears in the tweet text without including #; d(i) and d(j) are the total number of tweets in the set that contain i and j at least once; D is the total number of tweets in the social network event; (C) Establishment of relationship weights between social network event tags. The implicit semantic relationship only establishes associations between tag pairs with positive PMI values; finally, the relationship weights between social network event tags are determined by the summation method: R(i,j) = {R ex (i,j) + R im (i,j)} For social network events, the above process establishes an event label relationship graph G = <V, E> in the data of the observable time window; (D) Sample a sub-network \(G\) of a fixed size for the target label \(i\) from the global network event label relationship graph \(G\). i , rather than directly processing the global network \(G\) itself with a lot of noisy information. The sampling method in this step selects the restart random walk algorithm RWR. The RWR method is used to select the relevant labels that are most likely to affect the propagation, and the upper limit of the associated label nodes is constrained, that is: G i = RWR(G, i, m, R) Where G represents the global relationship graph of event tags, m represents the constraint number of sampled nodes, i represents the target tag, and R represents the weight R(i,j) between the nodes of the event label relationship graph.

4. A method for predicting the popularity of network event tags based on multi-tag influence according to claim 2, characterized in that, S302 specifically includes: (A) Extraction of label semantic features; Supplement the semantic information of the label in the tweet where the label is located within the observable time window, obtain the tweet s with the label that has the largest number of retweets within the observable time, and use it as the feature to explain the label semantics; (B) Embedded representation of the semantic features of event labels. After establishing the text library for the semantic interpretation of labels, the sentenceTransformer5 interface based on the bert model is called to perform the initial semantic vector embedding of the labels, formalized as: H s = {bert2sentence(i)} where \(0\leq i\leq|V|\), \(H\) s \(\in\mathbb{R}\) |v|×d , \(|V|\) represents the number of all labeled nodes, and \(d\) represents the embedding dimension; (C) Static semantic relationship attribute graph construction, using the event label propagation local influence relationship graph G in step 301 i Construct a static semantic relationship attribute graph with the node attributes obtained in the previous step: Among them, V i is the set of associated nodes of the target label i, E i is the set of label-node relationship, is the set of semantic feature representations of node V i and 5. A method for predicting the popularity of network event tags based on multi-tag influence according to claim 1, characterized in that, S303 specifically includes: (A) Calculate the social influence of the participating group. Define the social influence of label i within window t. For each label node, perform the following calculations: where N t (n) represents the total number of labeled tweets in the time window t, N t (E) represents the total number of tweets of all associated label nodes in the subgraph under the time window t. This metric expresses the overall influence degree of the group participating in this label on the event label at time t; (B) Calculate the propagation influence of the participating group. Define the group propagation influence of hashtagi at time t; for each node, perform the following calculations: Among them, represents the total number of followers of the users who post tweets with the label at time window t; represents the total number of users who post event-related tweets under time window t, and this indicator expresses the degree of attention of the group participating in the event label to the current label group; (C) Construct label dynamic group attribute H inf (t), select the decisive dynamic group influence feature H for propagation inf (t) as the node attribute; use two key indicators of label influence as the influence evaluation criteria for the group participated by the label under time series, and obtain the dynamic propagation attribute H inf (t), formalized as follows: H inf (t) = [O(t), M(t)] (D) Construction timing diagram sequence By dynamically propagating attribute H inf (t) Construct an attribute graph G based on dynamic attributes and dynamic relationships t , and then obtain the event label propagation timing attribute graph sequence within each time window: Among them Then, for each label i of the hotspot events in the dataset, the feature map of the target label i was constructed according to the above method as the sample input of the deep learning model 6. A method for predicting the popularity of network event tags based on multi-tag influence according to claim 1, characterized in that, Step 4 specifically includes: S401. Learn the propagation relationship feature embedding layer. The propagation relationship feature embedding is responsible for learning the propagation influence relationship between different labels, including the static semantic feature aggregation representation and the dynamic propagation feature aggregation representation; select the graph neural network GAT that can learn the influence relationship weights as the method for graph neural network feature representation, and learn different weight coefficients between nodes during the supervised learning process of GAT, and obtain the hidden mutual influence relationship in the node representation; The input of GAT consists of two parts, the node feature vector H ∈ R |V|×d and the adjacency matrix N t ∈ R |V|×|V| , where d represents the dimension of the features and n represents the number of nodes in the subgraph; for each node v i in the graph and the attribute h i , there is: H = [h1, h2,..., h1,] T H o = [h′1, h′2,..., h′ n ,] T , H o ∈ R |V|×d′ where α ∈ R 2×d′ , Θ represents a trainable weight matrix, j ∈ N(i) indicates that there is an edge between label node j and i in the adjacency matrix, that is, it means j and i are related; the LeakyReLu method is selected as the non-linear activation function, a i,j represents the weight of the mutual influence relationship between label node i and label node j, H O is a matrix composed of the output embedding vectors, |V| represents the number of graph nodes, d’ represents the dimension of the output node features, and || represents the concatenation operation; For the static semantic relationship attribute graph, is used as the input to the learning propagation relationship feature embedding layer to obtain the static semantic node vector representation matrix H sem , For the dynamic propagation time series attribute graph sequence, for each propagation attribute graph under each time window is used as the input to obtain the tensor H' composed of t windows within the observable time dym : Where t represents the number of time windows, |V| represents the number of subgraph nodes, and F2′ represents the feature dimension of the output nodes; S402. The feature fusion module is responsible for fusing the semantic features and propagation features of the label nodes, and then using them as the input for the next layer. The model will use H sem to broadcast in the matrix H' dym That is, the feature representation layer is formalized as follows: H f = H sem || H′ dym || represents a concatenation operation, so we have S403. According to the characteristic that there is local strong correlation in label semantics during the development of the event situation, the model applies a routing mechanism to vote on the group effect of label nodes, captures the local-to-global relationship in the graph, and then infers the local-to-global hierarchical relationship through multiple rounds of iteration, and finally obtains the temporal graph feature representation.

7. A method for predicting the popularity of network event tags based on multi-tag influence according to claim 6, characterized in that The process of mapping from the timing attribute graph in S403 to the graph embedding vector h G is divided into three steps in this step: (A) Hierarchicalize the temporal graph, and establish the relationship between the low-level local and high-level global of the temporal graph through the voting mechanism; it is stipulated that v represents the capsule node in the low-level voting graph, and u represents the capsule node in the high-level routing graph. First, use the feature fusion layer and the temporal graph as the initial voting matrix, that is: Among which N t represents the adjacency matrix of the low-level voting graph, represents the voting representation vector of the initial low-level network in the graph capsule network, and each vector is regarded as v i is the voting weight representation for the capsule node u j in the high-level cluster, |N| represents the number of nodes in the capsule, and F3′ = F1′ + F1′ represents the dimension of the input vector; (B) Establish dynamic routing. The task of this process is to iteratively calculate the routing weight C between the lower-layer node v and the higher-layer capsule u i,j , that is, which lower-layer v j groups can activate the higher-layer cluster node u j , to obtain the local-to-global activation relationship; perform a weighted calculation on the voting in step (A) to obtain the routing weight C from v j in the lower-layer graph to the locally aggregated higher-layer graph i,j j, that is: Initialize it to 0. Then calculate the dynamic routing weights through R iterations using the following three formulas. The three formulas are as follows: b i,j = b i,j + v j|i .u j Among them, the role of the squash non-linear "squeezing" function is to calculate the capsule u j for the routing possibility of the node v j|i in it, that is, the probability that the node v i|j votes for u j where v i|j ·u j is used to calculate the consistency between each group of votes and the high-level capsule; after R iterations, this process obtains the capsule node u of the high-level aggregated graph and the adjacency matrix of the high-level abstraction, expressed as: G route = (A, u) A = C T N, C ∈ R |V|×|U| where N is the adjacency matrix in the low-level voting graph, |V| represents the number of nodes in the low-level voting graph, |U| represents the number of nodes in the high-level routing graph, C represents the routing weight matrix from the low level to the high level formed by C i,j constituting, C ∈ R |V|×|U| , A is the adjacency matrix of the high-level routing graph, u is the eigenvector of the nodes in the high-level routing graph, and step B is defined as the following transformation: U,A=Route(Vote(V,N)) That is: U,A=RV(V,N) (C) Establish the temporal attribute graph representation, and then repeat the above steps (A) and (B) to abstract the high-level cluster graph into the representation of the whole graph embedding again. On the basis of the graph capsule network, retain the features of the influence of local hashtag labels on propagation, that is: The 1 in the formula represents the graph representation that only aggregates one node after being abstracted to a higher level. The feature vector representation of this node is the propagation influence representation vector of the label relationship graph under the current time window t; for each temporal attribute graph, use the local aggregation layer to perform the graph representation process:

8. A method for predicting the popularity of network event tags based on multi-label influence according to claim 1, characterized in that S5 specifically includes: S501. Apply the long short-term memory (LSTM) kernel through the representation vector h of different temporal subgraphs in the sample t , and the specific calculation formula is as follows: h t =tanh(c i )*o i where h t is the implicit feature output at the t-th moment, denotes the Hadamard product, U j , W j , b j , j ∈ ({z, f, o, c}) are learnable parameters, z i , f i and o i are the forget gate vector, input gate vector, and output gate vector of the t-th window feature respectively; Finally, predict the result at time t + 1 through the fully connected layer: Δy′ = σ(Wh t ) S502. Obtain the true value label of the sample, for the local propagation influence feature map G of the network event label i i , count the of this sample within the t+1 snapshot and set this value as the true value label y of the sample i ; Using the sum of the number of retweets and the number of original posts in the form of tags as the attention index of the user group, that is where represents how many people have retweeted the tweets with event tags, represents how many original tweets cover the target tags; S503. For the prediction of network popularity, regard the popularity prediction as a regression model, and use MLSE as the objective loss function based on the regression model: where Δy′ represents the predicted popularity metric and y i represents the actual spread metric.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the network event label popularity prediction method based on multi-label influence as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, it implements the method for predicting the popularity of network event tags based on multi-tag influence as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Social network-oriented hot event prediction method

    CN113806534A

  • Social network influence prediction method and device based on heterogeneous network

    CN114090902A