An unsupervised social media summarization method based on denoising graph autoencoder

Through an unsupervised method based on denoising graph autoencoder, using social relationship networks and residual graph attention network encoder, the problem of relying on labeled data and difficulty in using social relationship information in the prior art is solved, and high-quality social media digest generation is achieved.

CN115017299BActive Publication Date: 2025-05-13TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210393787.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-15
Publication Date
2025-05-13
Estimated Expiration
2042-04-15

AI Technical Summary

Technical Problem

Existing social media digest methods rely on labeled data, which is costly to build, and traditional methods are difficult to effectively utilize social relationship information, resulting in low quality of digests.

Method used

Using an unsupervised social media digest method based on denoising graph autoencoder, a post-level social relationship network is constructed, and a residual graph attention network encoder and a sparse reconstruction digest extractor are introduced to automatically identify and remove noise relationships to generate high-quality digests.

Benefits of technology

Without labeled data, the quality and efficiency of the digest are improved, the impact of noise relationships is reduced, and the performance of ROUGE evaluation indicators is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115017299B_ABST
    Figure CN115017299B_ABST
Patent Text Reader

Abstract

The present invention discloses an unsupervised social media summarization method based on a denoising graph autoencoder. According to sociological theory, a post-level social relationship network is constructed, and a pre-trained BERT model is used to obtain content encoding of the post as the initial content representation of the post; two types of noise relations are defined and corresponding noise functions are set to construct a pseudo social relationship network with noise relations; the sampled pseudo social relationship network instance and the initial content representation of the post are simultaneously used as inputs of a residual graph attention network encoder, and the residual graph attention network encoder encodes the post according to the initial content representation and social relationship of the post to obtain a vector representation of the post; a decoder is constructed, and the residual graph attention network encoder and the decoder together constitute a denoising graph autoencoder structure, and the denoising graph autoencoder can learn to remove noise relations in the post-level social relationship network, and finally obtain an accurate post representation; a summary extractor based on sparse reconstruction is used to select the final summary.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing and social media data mining, and in particular to an unsupervised social media summarization method based on a denoising graph autoencoder. Background Art

[0002] With the development and popularization of Internet technology, social media platforms have gradually become a new type of information production and dissemination medium, and have gradually occupied an increasingly important position in various aspects of social production and life. However, the rapid increase in the volume of content on social media has led to a serious information overload problem, posing a more severe challenge to the efficient retrieval of information. Ordinary users often find it difficult to search and obtain effective and interesting information in a massive amount of noisy information, which seriously reduces the efficiency of information retrieval.

[0003] Automatic summarization of social media text aims to automatically generate concise summary descriptions for a large collection of social media content. This technology can effectively alleviate the problem of information overload on social media and help improve the efficiency of users' retrieval of effective information. The main summarization methods at present can be generally divided into two modes: extractive summarization and generative summarization. Among them, extractive summarization mainly selects the most representative text units (words, sentences or fragments) with large information volume, low redundancy and wide coverage from the input original text to form the final summary; generative summarization methods involve the text generation process, which generates corresponding summary descriptions by understanding the semantics of the original input text and using text generation technology. In recent years, due to the development of many new technologies such as sequence-to-sequence framework (Seq2Seq), Transformer model, contrastive learning and large-scale pre-training models, both extractive and generative automatic text summarization technologies have made significant progress.

[0004] However, existing methods usually rely on large-scale annotated paired training data (i.e., text-summary pairs). Currently, the acquisition of such annotated training data usually requires manual data annotation, which is very costly and cannot be used in large-scale training scenarios. In the field of social media, the construction of annotated data is even more difficult: on the one hand, when annotators summarize the content of a specific topic, they need to read all posts related to the topic and then write corresponding summaries for these posts. However, the number of posts on social media is too large, and manual reading will incur unbearable labor costs; on the other hand, since the content on social media is highly real-time and topic-sensitive, the results of annotation under a specific topic cannot be applied to other topic areas. Therefore, data annotation work needs to be performed under each topic, which will consume a lot of manpower and material resources. Furthermore, when traditional text summarization methods are migrated to social media data, due to the fact that the characteristics of text on social media are quite different from those of traditional long documents, such as shorter text length, diverse expressions, and informality, traditional summarization methods usually find it difficult to achieve satisfactory results.

[0005] Existing social media summarization research mainly extracts features from each post independently based on the content of the post, and then uses algorithms such as graph sorting or clustering to extract posts with higher importance as summaries. These methods have certain disadvantages: (1) Since the length of posts on social media is usually short, the content of a single post often contains incomplete or ambiguous information and cannot provide sufficient information, resulting in sparse and inaccurate features of the post; (2) Social media relies on users to actively spread and receive information through social contact, which effectively promotes the spread of information. Therefore, posts on social media are embedded in a social network structure rather than being independent of each other. Previous methods only focus on the text content features of the post and ignore the social structure features of the post, resulting in the loss of social relationship information of the post.

[0006] Some works have attempted to facilitate the analysis of content on social media by utilizing some simple social signals provided on social media, such as the number of followers of the author, the number of reposts and likes of a post, etc. Further work has verified the impact of social relationships on the relevance of content in social networks from the perspective of sociological theory, and proposed that in a short period of time, posts with social relationships are more likely to contain similar content and opinions. This sociological theory points out the relationship between social relationships and text content from a macro perspective. However, at a micro level, there are often noise relationships that do not conform to this sociological theory. Specifically, there are two situations: (1) Two posts have a social relationship, but have a low correlation in content. This noise relationship is defined as a false relationship; (2) There is no direct social relationship between two posts, but they have a high correlation in content. This noise relationship is defined as a potential relationship. The existence of these two types of relationships poses further challenges to the effective use of social relationships. Summary of the invention

[0007] The purpose of the present invention is to overcome the deficiencies in the prior art and to provide an unsupervised extractive social media summarization method that is more robust to noise relations.

[0008] The objective of the present invention is achieved through the following technical solutions:

[0009] An unsupervised social media summarization method based on denoising graph autoencoder, comprising the following steps:

[0010] S1. Construct a post-level social relationship network based on sociological theory, define the noise-free relationship in the post-level social relationship network, that is, the real social relationship network, and use the pre-trained BERT model to obtain the content encoding of the post as the initial content representation of the post;

[0011] S2. According to the user's social behavior and habits, two types of noise relations, false relations and potential relations, are defined; by setting the corresponding noise function, instances of false relations and potential relations are added to the original post-level social relation network, and a post-level social relation network with noise relations is constructed, which is called a pseudo-social relation network; several pseudo-social relation networks generated are sampled, and the sampled pseudo-social relation network instances and the initial content representation of the post are used as the input of the residual graph attention network encoder at the same time. The residual graph attention network encoder includes a multi-head attention mechanism. The residual graph attention network encoder encodes the post according to the initial content representation and social relations of the post to obtain a vector representation of the post;

[0012] S3. Construct a decoder. The decoder and the residual graph attention network encoder together form a denoising graph autoencoder. The decoder reconstructs the real social relationship network based on the vector representation of the post to capture the social relationship information between posts. On the other hand, it reconstructs the semantic relationship between the post and the words it contains to capture the text content information of the post. At the same time, since the real social relationship network without noise relationship is reconstructed, the residual graph attention network encoder and decoder can learn to exclude the noise relationship in the post-level social relationship network, and finally obtain accurate post representation.

[0013] S4. According to the post representation obtained in step S3, a summary extractor based on sparse reconstruction is used to select the final summary, and the post with the highest reconstruction coefficient is iteratively selected to add to the final summary set. This process is repeated until the length limit of the summary is reached.

[0014] Furthermore, step S1 is specifically as follows: the post-level social relationship network is composed of a node set and an edge set, each node in the node set represents a post, and each edge in the edge set represents the social relationship between corresponding posts; there are two types of social relationships between posts, namely expression consistency relationship and expression contagiousness relationship; expression consistency relationship refers to the relationship between posts posted by the same user, and when constructing the post-level social relationship network, an edge is established between post nodes with expression consistency relationship; expression contagiousness relationship refers to the relationship between posts posted by users with direct interactive relationships, wherein the direct interactive relationship refers to the interactive relationship of attention, forwarding, and commenting between users, and when constructing the post-level social relationship network, an edge is established between post nodes with expression contagiousness relationship.

[0015] (101) The post-level social relationship network is formally described as follows: represents a set of posts, N is the number of posts, and s i (1≤i≤N) represents the i-th post; Represents a user set, which contains a total of M users, where u i (1≤i≤M) represents the i-th user; for user u i ,make Represents user u i Neighbor users, that is, users u i A collection of users with direct social connections; Indicates all the i A collection of published posts; a post-level social relationship network is constructed according to the following rules in represents a set of nodes, each node corresponds to a post, ε represents a set of edges between nodes, and each edge corresponds to a set of social relations between posts; in expressing consistent social relations: if a post where u k represents the kth user, then the post s i With Posts j Create an edge between ij ∈ε; in expressing infectious social relations: if the post And there are or Then for posts i With s j Create an edge between ij ∈ε; Based on these two rules, a post-level social relationship network is constructed It only contains the post node collection And the relationship between nodes, that is, the edge set ε={e 11 ,e 12 ,…,e NN}; Post-level social relationship network constructed The corresponding adjacency matrix is ​​recorded as Among them A ij >0 means post node s i With s j There is a social relationship between them, otherwise A ij =0;

[0016] (102) The content encoding of the post is obtained using the pre-trained BERT model as the initial content representation of the post, as follows:

[0017] For each post i , input the post into the pre-trained BERT model, and then regard the representation of the sentence start symbol of the last layer of the pre-trained BERT model as the initial content representation of the post; as shown in formula (1):

[0018] x i =BERT(s i ) (1)

[0019] where x i Indicates posts i The initial content representation of all N posts is finally obtained as X = [x1,…,x N ].

[0020] Furthermore, step S2 specifically includes:

[0021] (201) The definitions of two types of noise relations, spurious relations and latent relations, are as follows:

[0022] In social networks, posts connected by social relationships usually have more similar content or opinions. However, in real life, most social networks on social media are pseudo-social relationship networks, which contain noisy relationships. Based on the observation of real social media data, two types of noisy relationships are defined:

[0023] (a) False relationship: If there is a social relationship between two posts, but their content relevance is less than a set threshold, the social relationship between the two posts is defined as a false relationship;

[0024] (b) Potential relationship: If there is no social relationship between two posts, but their content relevance is greater than a set threshold, then the two posts are defined as having a potential relationship;

[0025] Set the noise function corresponding to the false relationship to relationship insertion, and set the noise function of the potential relationship to relationship loss, as follows:

[0026] (c) Relationship insertion: For any two unconnected post nodes in the post-level social relationship network, randomly add an edge to connect the two nodes;

[0027] (d) Relationship loss: For any two connected post nodes in the post-level social relationship network, the edge between them is randomly removed;

[0028] By adding instances of noise relations to the real social relation network, a pseudo social relation network is constructed as training data.

[0029] (202) Encode the posts using a residual graph attention network encoder;

[0030] After obtaining the pseudo social relationship network and the initial content representation of the post, in order to model the social relationship between posts, the residual graph attention network encoder is used to encode the post according to the initial content representation of the post and the social relationship between the posts to integrate the social relationship information and text content information of the post; the residual graph attention network encoder can be regarded as an information propagation model, which learns the representation of nodes in the post-level social relationship network by aggregating the information of neighbor nodes that are connected to the node by edges, where neighbor nodes refer to nodes that are connected by edges in the post-level social relationship network. At the same time, compared with the traditional graph convolutional network encoder (Graph Convolutional Networks, GCN), the residual graph attention network encoder can assign different weights to different neighbor nodes of the same node, thereby increasing the attention weight of important neighbor nodes, reducing the weight of low-correlation neighbor nodes, and learning more accurate node representations;

[0031] Formally, the residual graph attention network encoder represents the initial content of the node Adjacency matrix corresponding to the post-level social relationship network As input, where D is the dimension of the node feature representation and N is the number of nodes. The propagation rule of the residual graph attention network encoder is shown in formula (2) and formula (3):

[0032]

[0033]

[0034] Among them, H (l) is the hidden representation of the residual graph attention network encoder at layer l, A is the adjacency matrix corresponding to the post-level social relationship network, and A ij Representative posts i With s j The social relationship weight between them, I is the unit matrix, is the adjacency matrix corresponding to the post-level social relationship network after adding attention weights, Indicates that the post s after increasing the attention weight i With s j The relationship weight between , σ(·) represents the nonlinear activation function; It's posts i With Posts j The attention score between them at layer l; W (l) With b (l) is the learning parameter of the residual graph attention network encoder at layer l; to further integrate the initial content representation of the post, the initial content representation of the post is X = [x1,…,x N ] as the input of the residual graph attention network encoder, that is, let H (0) =X; the attention weight is calculated using scaled dot product attention [1] , the ordinary attention mechanism is extended to a multi-head attention mechanism, by mapping the potential representation to K different subspaces, K represents the total number of attention heads in the multi-head attention mechanism, each subspace is called an attention head, and the attention weight is calculated separately in each subspace:

[0035]

[0036]

[0037] where h i With h j Respectively represent posts i With Posts j The vector representation encoded by the residual graph attention network encoder; and They represent the post s in the kth attention head respectively. i With s j The attention score between and the normalized attention weight; (·) T Indicates transpose operation; D h is the dimension of the implicit representation in the attention calculation process. The superscript (l) indicating the number of layers is omitted here, and the superscript head is used instead. k to indicate the kth attention head; and is the corresponding learning parameter in the kth attention head; K attention weights are obtained by calculation of formula (4) and formula (5); K represents the total number of attention heads, which is K in total; k represents the kth one. The maximum pooling operation is used to automatically select the strongest relationship in all subspaces as the true relationship between the two post nodes, and the attention weights in the K attention heads are unified into the final attention score:

[0038]

[0039] α ij Indicates posts i With s j The final attention weight between the two layers; the connections between the layers in the ordinary graph attention network are replaced with residual connections to form a residual graph attention network, so that the residual graph attention network can directly transmit the input information to the output layer. Therefore, the encoding rule of the residual graph attention network encoder is modified to the following form:

[0040]

[0041] Where f(·) is the mapping function, which is implemented by a feedforward neural network with a nonlinear activation function:

[0042] f(H (l) )=σ(W f H (l) +b f ) (8)

[0043] Where W f With b f is the corresponding learning parameter in the mapping function, σ(·) represents the nonlinear activation function; during the encoding process, the depth of the residual graph attention network encoder Determines the distance of information propagation in the post-level social relationship network. The residual graph attention network encoder encodes the post according to the encoding rules of formula (7) and formula (8). The output of its last layer is is the vector representation of the encoded post, where represents the post s encoded by the residual graph attention network encoderi The vector representation of ;

[0044] Furthermore, step S3 is specifically as follows:

[0045] (301) Reconstruct the real social relationship network and the content of the posts using a reconstruction-based decoder.

[0046] In order to make the learned vector representation of posts contain both text content information and social relationship information between posts, a decoder based on two reconstruction objectives is designed. On the one hand, the decoder reconstructs the real social relationship network without noise relations to capture the social relationship information between posts, and on the other hand, reconstructs the text content contained in the posts, thereby capturing the text content information of the posts and enriching the vector representation of the posts.

[0047] For the reconstruction of the real social relationship network, the decoder predicts whether there is a social relationship between two post nodes based on the vector representation of the two nodes. Specifically, the inner product of the vector representation between the two nodes is used to predict the probability of the existence of a social relationship between the two nodes:

[0048]

[0049] in(·) T represents the transposition operation of the vector representation; for each pair of posts s i With s j , the decoder predicts the probability that there is a social relationship between them, where is the adjacency matrix corresponding to the post-level social relationship network output by the decoder, represents the post s predicted by the decoder i With s j The probability that there is a social relationship between them, and Respectively represent posts i With Posts j The vector representation obtained by encoding the residual graph attention network encoder, σ(·) represents the nonlinear activation function;

[0050] For text content reconstruction, we propose to reconstruct the relationship between posts and words, and retain the text content information of posts by reconstructing the words contained in each post. Since each post usually contains several words, the text content reconstruction process is modeled as a multi-label classification task:

[0051]

[0052] in and is the learning parameter of the decoder, V represents the vocabulary size; is the prediction result of the decoder, where Indicates posts i Contains the word w j probability;

[0053] Corresponding loss functions are designed for the above two reconstruction objectives. The overall training objective includes two parts. The first part of the loss is the loss of reconstructing the real social relationship network, denoted as L g , calculate the prediction results The binary cross entropy loss between the adjacency matrix A corresponding to the true post-level social relationship network:

[0054]

[0055] The second part of the loss function is the loss of reconstructing the post content, denoted by L c , calculate the decoder's prediction result With the real results i The binary cross entropy loss between:

[0056]

[0057] s ij is the real training label, indicating the post s i Does it contain the word w j If the post s i Contains the word w j , then s ij =1, otherwise s ij = 0. Finally, the two losses are combined using the balance parameter λ to obtain the final loss function L:

[0058] L=λL g +(1-λ)L c (13)

[0059] The residual graph attention network encoder and decoder are trained according to the loss function. After the training, an accurate post representation H = [h1,h3,…,h N ];

[0060] Furthermore, step S4 specifically includes:

[0061] (401) A sparse reconstruction based summary extractor is used to extract summaries based on the vector representation of the post.

[0062] In order to extract representative and important posts as the final summary, a summary extractor based on sparse reconstruction is used to extract the summary. Formally, given the residual graph attention network encoder, the accurate post representation H = [h1,h2,…,h N ], the abstract extraction process is modeled as a sparse reconstruction process:

[0063]

[0064] where ||·|| represents the Frobenius norm, is the reconstruction coefficient matrix, where each element V i,j Indicates posts j For refactoring posts i To prevent the extraction of duplicate and redundant content, a similarity matrix is ​​introduced To remove redundant information, if the post s i With Posts j If the cosine similarity of is higher than a certain threshold η, then otherwise ° represents the Hadamard product operation; to prevent the post from reconstructing itself, the diagonal elements of the reconstruction coefficient matrix V are restricted to 0 during the reconstruction process; β and γ are hyperparameters that control the weights of the corresponding regularization terms; H is the accurate post representation; ||·|| 2,1 represents the L21 norm and is defined as follows:

[0065]

[0066] Adding L21 constraints to the reconstruction coefficient matrix V can make each row of the reconstruction coefficient matrix sparse, that is, most of the elements in each row of the reconstruction coefficient matrix are 0, which means that each post can only be reconstructed by a limited number of posts to limit the length of the summary; the final score of each post is defined as the sum of the contribution of the post to the reconstruction of all other posts:

[0067]

[0068] Where score(s i ) indicates posts i Finally, all posts are sorted according to their final scores, and the posts with the highest scores are iteratively selected to join the final summary set. This process is repeated until the length limit of the summary is reached.

[0069] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:

[0070] 1. The present invention can extract summaries without labeled data. By introducing the social relationship information between posts, the social relationship features between posts are captured to make up for the sparse features caused by the short content of a single post;

[0071] 2. This paper proposes a denoising graph autoencoder structure that can automatically identify and remove unreliable noise relationships in social relationship networks without labeled data, alleviate the errors caused by noise relationships, and thus improve the reliability and accuracy of post representation. After learning the accurate representation of the post, the sparse reconstruction technology framework is used to identify the importance and redundancy of each post, and finally extract posts with higher importance and lower redundancy to form the final summary.

[0072] 3. Compared with the existing summary model, the social media text content summary obtained by the present invention improves the performance on the ROUGE evaluation index. At the same time, according to the experimental results, the denoising model can effectively reduce the ratio of noise relations in social networks, improve the network structure, and improve the accuracy of the summary.

[0073] 4. Compared with the traditional graph convolutional network encoder, the residual graph attention network encoder used in the present invention can assign different weights to different neighbor nodes of the same node, thereby increasing the attention weights of important neighbor nodes, reducing the weights of irrelevant neighbor nodes, and learning more accurate post representations; since the post vector representation obtained by the residual graph attention network encoder contains both text content information and social relationship information, the summary extraction process can identify the importance and novelty of the post from the two perspectives of content and social relationship, thereby generating a higher quality summary.

[0074] 5. Since different attention heads capture the relationship information between nodes from different spaces, the relationship between two nodes in different attention heads may be quite different. Therefore, the present invention adopts the maximum pooling operation to automatically select the strongest relationship in all subspaces as the true relationship between the two nodes, and unifies the attention weights in the K attention heads into the final attention score.

[0075] 6. Ordinary graph attention networks often have the problem of over-smoothing, and the present invention further replaces the connections between layers in the ordinary graph attention network with residual connections to form a residual graph attention network, so that it can directly transmit input information to the output layer.

[0076] 7. The present invention designs a decoder based on two reconstruction objectives, so that the learned vector representation of the post contains both the content information of the text and the social relationship information between the posts. Since the feature representation of the post contains both the content information and the social relationship information of the post, the summary process can identify the importance and novelty of the post from both the text content and the social structure, thereby generating summary content with large amount of information, high diversity and wide coverage. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 Schematic diagram of the overall architecture of the unsupervised social media summarization method based on denoising graph autoencoder provided by the present invention.

[0078] Figure 2 The performance results achieved by the present invention in social networks under each topic in two datasets are shown.

[0079] Figure 3 The distribution of spurious relationships, latent relationships, and the sum of both relationships in the social network is shown.

[0080] Figure 4a and Figure 4b The effect of different denoising methods and different noise ratios in the denoising function on the results is shown. DETAILED DESCRIPTION

[0081] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0082] The present invention provides an unsupervised social media summarization method based on a denoising graph autoencoder. Two major social media datasets are used to evaluate the performance of the method. The overall framework of the method is shown in the figure Figure 1 shown. Figure 1 The post-level social relationship network and post collection are the inputs of the model. The noise function adds noise relationship instances to the post-level social relationship network to obtain a pseudo social relationship network with noise relationships. The post collection is encoded by the BERT model to obtain the initial content representation of the post, which is input into the residual graph attention network encoder together with the pseudo social relationship network with noise for encoding, and finally outputs the post representation that combines social relationships and text content. The model of this paper is trained as a whole according to the two goals of real social relationship network reconstruction and text content reconstruction. After the training, the accurate post representation with the noise relationship removed is finally obtained. The accurate post representation finally learned is input into the summary extraction based on sparse reconstruction to extract the final summary.

[0083] (1) Construction of post-level social relationship network

[0084] This example selects data from mainstream social media network platforms at home and abroad for experimental verification, and selects Twitter.

[11] With Sina Weibo

[12] Two social media platforms were experimentally verified, and the release time of posts under each topic in the data was guaranteed to be within 5 days. For Twitter data, the main language of the text content in Twitter is English, and a total of 44,034 posts and 11,240 users are included, of which each user publishes at least one post, and each user has at least one social relationship. There are 4 standard reference summaries under each topic for evaluating the results. In the experiment, links, user names and other special characters in the posts were removed, and stem extraction and stop word removal were performed at the same time, and posts with a length of less than 3 words were filtered out. For Weibo data, Sina Weibo is one of the most popular social media platforms in China, so the data collected from Sina Weibo is used in the embodiment, including 130k posts and 126k users, including a total of 10 different topics, in which the posts are organized in a tree structure according to interactive relationships (such as replies and forwarding), and 3 standard reference summaries are provided under each topic for evaluating the results. The statistical details of the two data sets after preprocessing are shown in Table 1. In the experiment, the ROUGE evaluation standard is used, and the four evaluation results of ROUGE-1, ROUGE-2, ROUGE-L and ROUGE-SU* are mainly reported.

[0085] Formally, let represents a set of posts, N is the number of posts, and s i (1≤i≤N) represents the i-th post; Represents a user set, which contains a total of M users, where u i (1≤i≤M) represents the i-th user; for user u i ,make Represents user u i Neighbor users, that is, users u i A collection of users with direct social connections; Indicates all the i A collection of published posts; a post-level social relationship network is constructed according to the following rules in represents a set of nodes, each node represents a post, ε represents a set of edges between nodes, and each edge represents a social relationship between posts; in expressing consistent social relationships: if a post where u k represents the kth user, then the post s i With Posts j Create an edge between ij∈ε; in expressing infectious social relations: if the post And there are or Then for posts i With s j Create an edge between ij ∈ε; Based on these two rules, a post-level social relationship network is constructed It only contains the node collection And the relationship between nodes, that is, the edge set ε={e 11 ,e 12 ,…,e NN}; Post-level social relationship network constructed The corresponding adjacency matrix is ​​recorded as Among them A ij >0 means post node s i With s j There is a social relationship between them, otherwise A ij =0;

[0086] Table 1. Social media dataset details

[0087]

[0088] (2) Noise distribution observation

[0089] In order to analyze the distribution of noise relations in the constructed post-level social relationship network, this embodiment proposes a simple method for estimating the distribution of noise relations in the network. Generally speaking, if the content correlation of two posts with social relations is lower than the set threshold θ, they are considered to have a false relationship; if the content correlation of two posts without social relations is higher than the set threshold θ, they are considered to have a potential relationship. Here, the cosine similarity between the TFIDF representations of the posts is used as the content correlation between the posts. Let Represents the adjacency matrix of the constructed post-level social relationship network, where A ij >0 means post node s i With Posts j There is a social relationship between them, otherwise A ij = 0. For each node pair (s i ,s j ), if they are connected by social relationships (i.e. A ij >0) and the content correlation between them Φ ij If the value is lower than the threshold θ, then the post s is considered i With s j There is a false relationship between them; if the post s i With s jThere is no social connection between them (i.e. A ij =0) and the correlation between their contents Φ ij If the value is higher than the threshold value θ, then it is considered that there is a potential relationship between them. In the experiment, the TFIDF representation of the post is calculated, and the cosine similarity of the TFIDF representation between the two posts is calculated as the correlation in the content between the posts. The greater the cosine similarity of the TFIDF representation of the two posts, the higher the correlation in the content of the two posts is considered. Conversely, the smaller the cosine similarity of the TFIDF representation of the two posts, the lower the correlation in the content of the two posts is considered. This statistical result can preliminarily reflect the distribution of noise relationships in the post-level social relationship network. Since the strength of social relationships in different social relationship networks may vary greatly, the average value of all social relationships is used as the value of the threshold θ in the embodiment:

[0090]

[0091] where Φ ij It's posts i With Posts j The final statistical results for the Twitter and Weibo datasets are shown in Table 2. The results show the average proportion of noise relationships in the post-level social relationship network under all topics. The average noise proportion in the table includes false relationships and potential relationships.

[0092] Table 2. Statistics of the distribution of noisy relationships in the social relationship networks in the Twitter and Weibo datasets

[0093] Dataset False relationship ratio Potential relationship ratio Average noise ratio Twitter Dataset 38.61% 55.79% 55.37% Weibo Dataset 83.17% 52.66% 52.67%

[0094] (3) Denoising Image Autoencoder

[0095] First, use the pre-trained BERT model to extract features from the post text content. The process is as follows:

[0096] x i =BERT(s i ) (1)

[0097] where s i represents the i-th post, x i This is the initial content representation of the post. All N posts are encoded and finally the matrix is ​​obtained. Where D is the dimension of each post feature vector. For the input post-level social relationship network Using noise function to improve post-level social relationship network Add noise relationship instances to construct a pseudo social relationship network Real social network And the corresponding pseudo social relationship network Forming paired training data After obtaining the initial content representation of the post and constructing the pseudo social relationship network, the initial content representation X of the post is combined with the pseudo social relationship network Together they serve as the input of the residual graph attention network encoder. Formally, the residual graph attention network encoder propagates information according to the following rules:

[0098]

[0099]

[0100] Among them, H (l) is the hidden representation of the residual graph attention network encoder at layer l, A is the adjacency matrix corresponding to the post-level social relationship network, and A ij Representative posts i With s j The social relationship weight between them, I is the unit matrix, is the adjacency matrix corresponding to the post-level social relationship network after adding attention weights, Indicates that the post s after increasing the attention weight i With s j The relationship weight between , σ(·) represents the nonlinear activation function; It's posts i With Posts j The attention score between them at layer l; W (l) With b (l) is the learning parameter of the residual graph attention network encoder at layer l; to further integrate the initial content representation of the post, the initial content representation of the post is X = [x1,…,x N ] as the input of the residual graph attention network encoder, that is, let H (0) =X; the attention weight is calculated using scaled dot product attention [1] , the ordinary attention mechanism is extended to a multi-head attention mechanism, by mapping the potential representation to K different subspaces, each subspace is called an attention head, and the attention weight is calculated separately in each subspace:

[0101]

[0102]

[0103] where h i With h j Respectively represent posts i With Posts j The vector representation encoded by the residual graph attention network encoder; and They represent the post s in the kth attention head respectively. i With s j The attention score between and the normalized attention weight; (·) T Indicates transpose operation; D h is the dimension of the implicit representation in the attention calculation process. The superscript (l) indicating the number of layers is omitted here, and the superscript head is used instead. k to indicate the kth attention head; and is the corresponding learning parameter in the kth attention head; K attention weights are obtained by calculating formula (4) and formula (5); in this embodiment, K represents the total number of attention heads, which is K in total; k represents the kth one. The maximum pooling operation is used to automatically select the strongest relationship in all subspaces as the true relationship between the two post nodes, and the attention weights in the K attention heads are unified into the final attention score:

[0104]

[0105] α ij Indicates posts i With s j The final attention weight between the two layers; the connections between the layers in the ordinary graph attention network are replaced with residual connections to form a residual graph attention network, so that the residual graph attention network can directly transmit the input information to the output layer. Therefore, the encoding rule of the residual graph attention network encoder is modified to the following form:

[0106]

[0107] Where f(·) is the mapping function, which is implemented by a feedforward neural network with a nonlinear activation function:

[0108] f(H (l) )=σ(W f H (l) +b f ) (8)

[0109] Where W f With b f is the corresponding learning parameter in the mapping function, σ(·) represents the nonlinear activation function; during the encoding process, the depth of the residual graph attention network encoder Determines the distance of information propagation in the post-level social relationship network. The residual graph attention network encoder encodes the post according to the encoding rules of formula (7) and formula (8). The output of its last layer is is the vector representation of the encoded post, where represents the post s encoded by the residual graph attention network encoder i The vector representation of ;

[0110] After the residual graph attention network encoder obtains the vector representation of the post, the decoder decodes the vector representation of the post and reconstructs the real social relationship network without noise relations, so as to learn to identify and remove noise relations in the pseudo social relationship network; during the training process, the model as a whole is trained according to the following loss function:

[0111]

[0112]

[0113] in, represents the post s predicted by the decoder i With s j The probability of a social relationship between them, A ij For posts i With Posts j The actual social relationship between them; is the prediction result of the decoder, where Indicates posts i Contains the word w j The probability of ij is the real training label, indicating the post s i Does it contain the word w j If the post s i Contains the word w j , then s ij =1, otherwise s ij =0. Where L g is the loss of reconstructing the original network structure, L c is the loss of reconstructing the original post content, and finally the two parts of the loss are combined using the balance parameter λ to obtain the final loss L:

[0114] L=λL g +(1-λ)L c (11)

[0115] The model of this paper is trained as a whole according to the loss function. After the training, in the test phase, the real social relationship network and the initial content representation of the post are used as input. The input is encoded through the residual graph attention network encoder to obtain an accurate post vector representation H = [h1,h2,…,h N ];

[0116] (4) Summary extraction based on sparse reconstruction

[0117] After getting the accurate post vector representation H = [h1,h2,…,h N ]After that, a sparse reconstruction-based framework is used to identify the importance of posts. The reconstruction process is modeled as follows:

[0118]

[0119] The symbols are as described above. The final importance score of each post is calculated as follows:

[0120]

[0121] Then, the posts are sorted according to their importance scores, and the posts with the highest scores are iteratively extracted and added to the candidate summary set until the length limit of the summary is reached.

[0122] In the specific implementation process, various hyperparameters are set in advance. The representation dimension D of the post is set to 768. Since the probability distribution of the two noises is often different in different social networks, the probabilities of relationship insertion and relationship loss in the noise function are set to 0.4 and 0.1 for Twitter data; for Weibo data, the probabilities of the two noise functions are both set to 0.3. The balance parameter λ of the two parts of the loss in the final loss function is set to 0.8. The hyperparameter β=γ=1 in the summary selection stage, and the redundant term threshold θ is set to 0.1.

[0123] In order to verify the effectiveness of the method of the present invention, the method of the present invention (DSNSum) is compared with two methods. The first method only uses the text content in social media to extract summaries, specifically including:

[0124] Centroid [2] Centrality-based features are used to identify sentences that are highly close to the cluster center as summaries.

[0125] LSA [3] The feature matrix is ​​decomposed by using SVD technology, and the importance of the post is identified based on the size of the singular value after matrix decomposition.

[0126] Lexrank [4] It is a graph sorting algorithm similar to PageRank. It first builds a similarity network based on the similarity of the content between posts, then uses a graph sorting algorithm similar to PageRank in the similarity network to identify the importance of each post node, and extracts posts with higher importance as summaries.

[0127] DSDR [5] The summarization process is regarded as a reconstruction task, and the most representative posts are extracted as summaries by minimizing the reconstruction loss.

[0128] MDS-Sparse [6] Sparse coding-based technology is used to extract multi-document summaries, minimizing the loss of reconstructing the original documents under sparse constraints to ensure the conciseness and importance of the summaries.

[0129] PacSum [8] It is a graph-based extractive summarization method that uses BERT to extract sentence features and models the document as a directed graph structure, while taking into account the relative position information between sentences.

[0130] Spectral [9] A spectrum-based hypothesis is proposed, the concept of spectrum importance is defined, and sentences with higher importance are extracted as summaries according to the spectral importance of the sentences.

[0131] The second type of method not only uses the text content features of the post, but also introduces the social relationship information between posts, including the following methods:

[0132] SNSR [7] Based on sociological theory, the social relations between posts are modeled as a regular term and introduced into the sparse reconstruction framework, so as to provide additional guidance for the summary extraction process.

[0133] SCMGR

[10] A graph convolutional network is used on the social relationship network between posts to encode post representations that combine text content and social structure, and the learned fusion representation is input into the sparse reconstruction framework to extract important posts.

[0134] The evaluation index of experimental performance adopts the ROUGE evaluation standard, which includes four indicators: ROUGE-1, ROUGE-2, ROUGE-L, and ROUGE-SU*. ROUGE-N measures the overlap of N-grams between the output summary and the standard summary. In this experiment, ROUGE-1 and ROUGE-2 standards are used for evaluation; ROUGE-L measures the longest common subsequence between the output summary and the standard summary; ROUGE-SU* measures the matching degree of 1-gram words and 2-gram phrases between the output summary and the standard summary, while allowing discontinuity between words. In subsequent experiments, the above four indicators are recorded as R-1, R-2, RL and R-SU* respectively.

[0135] Table 3 shows the experimental results of the model and all the comparison methods on the two datasets. The higher the ROUGE score, the better the performance of the model. Tables 4 and 5 show the performance of the model on Twitter.

[11] And the degradation experimental results on Weibo data, where DSNSum is the experimental result of the complete model, w / o denoising means the performance after removing the denoising module; w / o GAT means the performance after removing the residual graph attention encoder.

[0136] Table 3 Performance of the proposed method and other methods on Twitter and Weibo datasets

[0137]

[0138] Table 4 Degradation experimental results of the proposed method on Twitter data

[0139] Twitter data R-1 R-2 RL R-SU* DSNSum 46.51 14.29 44.16 20.76 w / o denoising 45.02 13.33 42.72 19.83 w / o GAT 44.14 12.65 41.68 19.10

[0140] Table 5 Degradation experimental results of the proposed method on Weibo data

[0141] Weibo data R-1 R-2 RL R-SU* DSNSum 37.01 10.98 14.22 13.06 w / o denoising 35.31 9.76 13.43 12.07 w / o GAT 34.36 8.93 13.29 11.12

[0142] As shown in Table 3, the proposed method achieves the highest performance on Twitter data, surpassing all other comparison methods; on Weibo data, it is slightly lower than the SCMGR model under the RL standard, and surpasses all other comparison models in other standards. The experimental results prove the effectiveness of the proposed method. In the degradation experiment, it can be seen from Tables 4 and 5 that removing any module will lead to a decrease in performance, proving that each module has a certain promoting effect on the overall model. Among them, after removing the denoising module, the model performance has declined, which proves that the noise relationship will introduce additional noise information to the summary process, thereby damaging the summary result. The denoising module reduces the impact of the noise relationship on the summary by identifying and removing the noise relationship in the network, thereby improving the quality of the summary. In addition, after removing the graph attention network, the performance of the model has dropped significantly. This phenomenon shows that considering the social relationship information in the post-level social relationship network can effectively promote the analysis of content in the social media environment. On the one hand, the graph attention network can aggregate relevant background information from adjacent neighbor nodes in the post-level social relationship network to alleviate the problem of insufficient content of a single post. On the other hand, the topological structure characteristics of the post-level social relationship network can provide additional clues for the importance identification of posts from a sociological perspective.

[0143] In order to further analyze whether the denoising graph autoencoder module proposed in the method of the present invention has the effect of removing noise relations and improving the network structure, additional experimental verification was carried out. While keeping the post representation unchanged (the post representation encoded using the same pre-trained BERT model), the denoised network was used to calculate the proportion of noise relations in the network. The results are shown in Table 6.

[0144] Table 6. The percentage of noisy relationships in the network after denoising in Twitter and Weibo data. The values ​​in brackets indicate the decrease compared to before denoising.

[0145] Dataset False relationship rate Potential relationship rate Average noise ratio Twitter data 13.60%(↓25.01%) 54.93%(↓0.86%) 54.50%(↓0.87%) Weibo data 45.29%(↓37.88%) 49.48%(↓3.18%) 46.57%(↓6.10%)

[0146] From the results in the table, we can see that without changing the representation of the post text content, the overall noise ratio in the network has decreased after denoising, proving the effectiveness of the denoising process. After denoising, the proportion of false relationships in Twitter and Weibo data decreased by 25.01% and 37.88% respectively, which shows that the denoising module is better at removing false relationships in the network.

[0147] In order to verify whether the post representation learned by the denoising graph autoencoder (DGAE) is better than the original BERT representation, the distribution of noise relations in the network is compared when the DGAE representation learned by the method of the present invention is kept unchanged with the BERT representation. Since the value of the threshold θ when calculating the noise relationship will seriously affect the distribution of the noise relationship, the experiment shows the noise distribution under different θ values. Specifically, the calculation method of the threshold θ is shown in the following formula:

[0148] θ=minΦ+δ*(maxΦ-minΦ)

[0149] Where Φ is the semantic similarity matrix between posts, δ is the adjustment parameter, and the experimental results are as follows Figure 3 shown.

[0150] Depend on Figure 3 It can be seen that as the threshold θ increases, the potential relationship rate decreases and the false relationship rate increases. The total noise relationship rate is generally maintained at a high level. After DGAE denoising, the potential relationship rate drops significantly and the false relationship rate also shows a low level. Most importantly, the total noise relationship rate has a significant downward trend compared to before denoising, proving that the DGAE representation can effectively remove noise relationships in the network.

[0151] Figure 3 The x-axis represents the value of the threshold δ. Sub-graphs (a) and (c) correspond to the representations encoded using the BERT model, and sub-graphs (b) and (d) correspond to the representations learned using the denoising graph autoencoder model.

[0152] In order to analyze the effect of the order and ratio of the two noise relations in the noise function on the model performance, additional experiments were conducted. By adjusting the probability of the two noise relations in the noise function, the trend of the model performance was observed. In addition, the order of adding the two noise relations was further adjusted to see whether it had an effect on the performance of the model. The order of adding the two noise relations was recorded as insert first and then lose, and lose first and then insert, and the change in model performance was observed. The results are shown in Figure 2. Figure 4a and Figure 4b shown.

[0153] Figure 4a and Figure 4b The influence of different denoising methods and different noise relationship adding probabilities on the experimental results are shown. Figure 4a Indicates the situation where false relationships are added first and then potential relationships are added in the noise function. Figure 4b It indicates the case where the potential relationship is added first and then the false relationship is added. The horizontal axis represents the insertion probability in the noise relationship, and the vertical axis represents the loss probability in the noise relationship.

[0154] The above contents are intended to illustrate the technical solution of the present invention schematically, and the present invention is not limited to the embodiments described above. Without departing from the scope of the present invention and the scope protected by the claims, a person skilled in the art may make many specific changes under the guidance of the present invention, which are all within the protection scope of the present invention.

[0155] References:

[0156] [1]Vaswani A,Shazeer N,Parmar N,et al.Attention is All You Need[C].InProceedings of the 31st International Conference on Neural InformationProcessing Systems,2017:6000–6010.

[0157] [2]Dragomir Radev,Sasha Blair-Goldensohn,and ZhuZhang.2001.Experiments in Single and Multi-Document Summarization UsingMEAD.In First Document Understanding Conference.1-8

[0158] [3]Yihong Gong and Xin Liu.2001.Generic Text Summarization UsingRelevance Measure and Latent Semantic Analysis.In Proceedings of the 24thAnnual International ACM SIGIR Conference on Research and Development inInformation Retrieval.19–25

[0159] [4]Gunes Erkan and Dragomir Radev.2011.LexRank:Graph-based LexicalCentrality As Salience in Text Summarization.Journal of ArtifcialIntelligence Research 22(Sept.2011),457–479

[0160] [5]Z.He,C.Chen,J.Bu,C.Wang,L.Zhang,D.Cai,and X.He.2012.Documentsummarization based on data reconstruction.In Twenty-sixth AAAI Conference onArtifcial Intelligence.620–626

[0161] [6]He Liu,Hongliang Yu,and Zhi-Hong Deng.2015.Multi-DocumentSummarization Based on Two-Level Sparse Representation Model.In Proceedingsof the Twenty-Ninth AAAI Conference on Artifcial Intelligence.196–202

[0162] [7]Ruifang He and Xingyi Duan.2018.Twitter Summarization Based onSocial Network and Sparse Reconstruction.In Proceedings of the Thirty-SecondAAAI Conference on Artifcial Intelligence.5787–5794

[0163] [8]Hao Zheng and Mirella Lapata.2019.Sentence Centrality Revisitedfor Unsupervised Summarization.In Proceedings of the 57th Annual Meeting ofthe Association for Computational Linguistics.6236–6247

[0164] [9]Baobao Chang Kexiang Wang and Zhifang Sui.2020.A Spectral Methodfor Unsupervised Multi-Document Summarization.In Proceedings of the2020Conference on Empirical Methods in Natural Language Processing.435–445

[0165]

[10] Huanyu Liu,Ruifang He,Liangliang Zhao,Haocheng Wang,and RuifangWang.2021.SCMGR:Using Social Context and Multi-Granularity Relations forUnsupervised Social Summarization.In Proceedings of the 30 th ACM InternationalConference on Information and Knowledge Management.1058-1068

[0166]

[11] Ruifang He,Liangliang Zhao,and Huanyu Liu.2020.TWEETSUM:Eventoriented Social Summarization Dataset.In Proceedings of the 28thInternational Conference on Computational Linguistics.5731–5736

[0167]

[12] Jing Li, Wei Gao, Zhongyu Wei, Baolin Peng, and Kam-FaiWong.2015.Using Content-level Structures for Summarizing Microblog RepostTrees.In Proceedings of the 2015Conference on Empirical Methods in NaturalLanguage Processing.2168–2178.

[0168] The present invention is not limited to the embodiments described above. The above description of the specific embodiments is intended to describe and illustrate the technical solution of the present invention. The above specific embodiments are merely illustrative and not restrictive. Without departing from the scope of the present invention and the scope of protection of the claims, a person of ordinary skill in the art can also make many forms of specific changes under the guidance of the present invention, which all fall within the scope of protection of the present invention.

Claims

1. An unsupervised social media summarization method based on denoising graph autoencoder, characterized in that The following steps are involved: S1. Construct a post-level social relationship network based on sociological theory, define the noise-free relationship in the post-level social relationship network, that is, the real social relationship network, and use the pre-trained BERT model to obtain the content encoding of the post as the initial content representation of the post; S2. Based on the user's social behaviors and habits, two types of noise relationships are defined: false relationships and potential relationships; By setting the corresponding noise function, instances of false relationships and potential relationships are added to the original post-level social relationship network to construct a post-level social relationship network with noise relationships, which is called a pseudo social relationship network. Several pseudo social relationship networks generated are sampled, and the sampled pseudo social relationship network instances and the initial content representation of the post are used as inputs of the residual graph attention network encoder. The residual graph attention network encoder includes a multi-head attention mechanism. The residual graph attention network encoder encodes the post according to the initial content representation and social relationship of the post to obtain a vector representation of the post. The definitions of two types of noise relations, false relations and potential relations, are as follows: (a) False relationship: If there is a social relationship between two posts, but their content relevance is less than a set threshold, the social relationship between the two posts is defined as a false relationship; (b) Potential relationship: If there is no social relationship between two posts, but their content relevance is greater than a set threshold, then the two posts are defined as having a potential relationship; Set the noise function corresponding to the false relationship to relationship insertion, and set the noise function of the potential relationship to relationship loss, as follows: (c) Relationship insertion: For any two unconnected post nodes in the post-level social relationship network, randomly add an edge to connect the two nodes; (d) Relationship loss: For any two connected post nodes in the post-level social relationship network, the edge between them is randomly removed; By adding noise relationship instances to the real social relationship network, a pseudo social relationship network is constructed as training data S3, construct a decoder, which together with the residual graph attention network encoder constitutes a denoising image autoencoder; The decoder reconstructs the real social relationship network based on the vector representation of the post to capture the social relationship information between posts, and reconstructs the semantic relationship between the post and the words it contains to capture the text content information of the post; the reconstructed social relationship network is the real social relationship network without noise relations, and the residual graph attention network encoder and decoder can learn to exclude the noise relations in the post-level social relationship network, and finally obtain accurate post representation; S4. According to the post representation obtained in step S3, a summary extractor based on sparse reconstruction is used to select the final summary, and the posts with the highest scores are iteratively selected to be added to the final summary set. This process is repeated until the length limit of the summary is reached.

2. The unsupervised social media summarization method based on denoising graph autoencoder according to claim 1, characterized in that: Step S1 is as follows: the post-level social relationship network consists of a node set and an edge set, each node in the node set represents a post, and each edge in the edge set represents the social relationship between corresponding posts; there are two types of social relationships between posts, namely expression consistency relationship and expression contagion relationship; expression consistency relationship refers to the relationship between posts posted by the same user, and when constructing the post-level social relationship network, an edge is established between post nodes with expression consistency relationship; expression contagion relationship refers to the relationship between posts posted by users with direct interaction relationship, where direct interaction relationship refers to the interaction relationship of attention, forwarding, and commenting between users, and when constructing the post-level social relationship network, an edge is established between post nodes with expression contagion relationship.

3. The unsupervised social media summarization method based on denoising graph autoencoder according to claim 2, characterized in that: In step S1: (101) The post-level social relationship network is formally described as follows: represents a set of posts, N is the number of posts, and s i (1≤i≤N) represents the i-th post; Represents a user set, which contains a total of M users, where u i (1≤i≤M) represents the i-th user; for user u i ,make Represents user u i Neighbor users, that is, users u i A collection of users with direct social connections; Indicates all the i A collection of published posts; Construct a post-level social relationship network according to the following rules in represents a set of nodes, each node corresponds to a post, ε represents a set of edges between nodes, and each edge corresponds to a set of social relations between posts; in expressing consistent social relations: if a post where u k represents the kth user, then the post s i With Posts j Create an edge between ij ∈ε; in expressing infectious social relations: if the post And there is u i ∈ or Then for posts i With s j Create an edge between ij ∈ε; Based on these two rules, a post-level social relationship network is constructed It only contains the post node collection And the relationship between nodes, that is, the edge set ε={e 11 ,e 12 ,…,e NN }; Post-level social relationship network constructed The corresponding adjacency matrix is ​​recorded as Among them A ij >0 means post node s i With s j There is a social relationship between them, otherwise A ij =0; (102) The content encoding of the post is obtained using the pre-trained BERT model as the initial content representation of the post, as follows: For each post i , input the post into the pre-trained BERT model, and then regard the representation of the sentence start symbol of the last layer of the pre-trained BERT model as the initial content representation of the post; as shown in formula (1): x i =BERT(s i ) (1) where x i Indicates posts i The initial content representation of all N posts is finally obtained as X = [x1,…,x N ].

4. The unsupervised social media summarization method based on denoising graph autoencoder according to claim 1, characterized in that: In step S2, a residual graph attention network encoder is used to encode the post according to the initial content representation of the post and the social relationship between the posts, so as to integrate the text content information and social relationship information of the post; the residual graph attention network encoder is regarded as an information propagation model, which learns the representation of the node in the post-level social relationship network by aggregating the information of the neighboring nodes connected to the node by edges, where the neighboring nodes refer to the nodes connected by edges in the post-level social relationship network; specifically as follows: The residual graph attention network encoder represents the initial content of the node Adjacency matrix corresponding to the post-level social relationship network As input, where D is the dimension of node feature representation and N is the number of posts; the propagation rule of the residual graph attention network encoder is shown in formula (2) and formula (3): Among them, H (l) is the hidden representation of the residual graph attention network encoder at layer l, A is the adjacency matrix corresponding to the post-level social relationship network, and A ij Representative posts i With s j The social relationship weight between them, I is the unit matrix, is the adjacency matrix corresponding to the post-level social relationship network after adding attention weights, Indicates that the post s after increasing the attention weight i With s j The relationship weight between , σ(·) represents the nonlinear activation function; It's posts i With Posts j The attention score between them at layer l; W (l) With b (l) is the learning parameter of the residual graph attention network encoder at layer l; to further integrate the initial content representation of the post, the initial content representation of the post is X = [x1,…,x N ] as the input of the residual graph attention network encoder, that is, let H (0) =X; the attention weight is calculated using scaled dot product attention [1] , the ordinary attention mechanism is extended to a multi-head attention mechanism, by mapping the potential representation to K different subspaces, K represents the total number of heads of the multi-head attention mechanism, each subspace is called an attention head, and the attention weight is calculated separately in each subspace: where h i With h j Respectively represent posts i With Posts j The vector representation encoded by the residual graph attention network encoder; and They represent the post s in the kth attention head respectively. i With s j The attention score between and the normalized attention weight; (·) T Indicates transpose operation; D h is the dimension of the implicit representation in the attention calculation process; the superscript (l) indicating the number of layers is omitted here, and the superscript head is used k to indicate the kth attention head; and is the corresponding learning parameter in the kth attention head; K attention weights are obtained by calculating formula (4) and formula (5); The maximum pooling operation is used to automatically select the strongest relationship in all subspaces as the true relationship between two post nodes, and the attention weights in the K attention heads are unified into the final attention score: α ij Indicates posts i With s j The final attention weight between The connections between the layers in the ordinary graph attention network are replaced with residual connections to form a residual graph attention network encoder, so that the residual graph attention network encoder can directly transmit the input information to the output layer. Therefore, the encoding rule of the residual graph attention network encoder is modified to the following form: Where f(·) is the mapping function, which is implemented by a feedforward neural network with a nonlinear activation function: f(H (l) )=σ(W f H (l) +b f ) (8) Where W f With b f is the corresponding learning parameter in the mapping function, σ(·) represents the nonlinear activation function; during the encoding process, the depth of the residual graph attention network encoder Determines the distance of information propagation in the post-level social relationship network. The residual graph attention network encoder encodes the post according to the encoding rules of formula (7) and formula (8). The output of its last layer is is the vector representation of the encoded post, where represents the post s encoded by the residual graph attention network encoder i The vector representation of is used for the subsequent summary extraction process based on sparse reconstruction.

5. The unsupervised social media summarization method based on denoising graph autoencoder according to claim 1, characterized in that: Step S3 is as follows: By setting a decoder based on two reconstruction objectives; the decoder reconstructs the real social relationship network without noise to capture the social relationship information between posts, reconstructs the text content contained in the posts, thereby capturing the text content information of the posts, and further enriching the vector representation of the posts; For the reconstruction of the real social relationship network, the decoder predicts whether there is a social relationship between two post nodes based on the vector representation of the two nodes. Specifically, the inner product of the vector representation between the two nodes is used to predict the probability of the existence of a social relationship between the two nodes: in(·) T represents the transposition operation of the vector representation; for each pair of posts s i With s j , the decoder predicts the probability that there is a social relationship between them, where is the adjacency matrix corresponding to the post-level social relationship network output by the decoder, represents the post s predicted by the decoder i With s j The probability that there is a social relationship between them, and Respectively represent posts i With Posts j The vector representation obtained by encoding the residual graph attention network encoder, σ(·) represents the nonlinear activation function; For text content reconstruction, we propose to reconstruct the relationship between posts and words, and retain the text content information of posts by reconstructing the words contained in each post. Each post usually contains several words, and the text content reconstruction process is modeled as a multi-label classification task: in and is the learning parameter of the decoder, Z is the dimension of the vector representation of the post obtained by the encoder, and V represents the vocabulary size; is the prediction result of the decoder, where Indicates posts i Contains the word w j probability; Corresponding loss functions are designed for the above two reconstruction objectives. The overall training objective includes two parts. The first part of the loss is the loss of reconstructing the real social relationship network, denoted as L g , calculate the prediction results The binary cross entropy loss between the adjacency matrix A corresponding to the real social relationship network: The second part of the loss function is the loss of reconstructing the post content, denoted by L c , calculate the decoder's prediction result With the real results i The binary cross entropy loss between: s ij is the real training label, indicating the post s i Does it contain the word w j If the post s i Contains the word w j , then s ij =1, otherwise s ij =0; finally, the two parts of the loss are combined using the balance parameter λ to obtain the final loss function L: L=λL g +(1-λ)L c (13) The residual graph attention network encoder and decoder are trained according to the loss function. After the training, an accurate post representation H = [h1,h2,…,h N ].

6. The unsupervised social media summarization method based on denoising graph autoencoder according to claim 1, characterized in that: Step S4 is specifically as follows: Given the residual graph attention network encoder, the accurate post representation H = [h1,h2,…,h N ], the abstract extraction process is modeled as a sparse reconstruction process: where ||·|| represents the Frobenius norm, is the reconstruction coefficient matrix, where each element V i,j Indicates posts j For refactoring posts i To prevent the extraction of duplicate and redundant content, a similarity matrix is ​​introduced To remove redundant information, if the post s i With Posts j If the cosine similarity of is higher than a certain threshold η, then otherwise ° represents the Hadamard product operation; to prevent the post from reconstructing itself, the diagonal elements of the reconstruction coefficient matrix V are restricted to 0 during the reconstruction process; β and γ are hyperparameters that control the weights of the corresponding regularization terms; H is the accurate post representation; ||·|| 2,1 represents the L21 norm and is defined as follows: Adding L21 constraints to the reconstruction coefficient matrix V can make each row of the reconstruction coefficient matrix sparse, that is, most of the elements in each row of the reconstruction coefficient matrix are 0, which means that each post can only be reconstructed by a limited number of posts to limit the length of the summary; the final score of each post is defined as the sum of the contribution of the post to the reconstruction of all other posts: Where score(s i ) indicates posts i Finally, all posts are sorted according to their final scores, and the posts with the highest scores are iteratively selected to join the final summary set. This process is repeated until the length limit of the summary is reached.

Citation Information

Patent Citations

  • Social relation topic model based social network friend recommendation method

    CN105740342A

  • Social media sentiment classification method and device based on knowledge graph

    CN111538835A