A keyword generation method and device integrating syntactic structure information

By introducing graph convolutional network and syntactic structure information into the keyword generation model, the shortcomings of the Seq2Seq model in capturing the global information of the article and modeling the "one-to-many" relationship are solved, and higher quality and diverse keyword generation are achieved.

CN114692605BActive Publication Date: 2025-05-06SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210415569.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-20
Publication Date
2025-05-06
Estimated Expiration
2042-04-20

AI Technical Summary

Technical Problem

The existing keyword generation model is based on the Seq2Seq model. The RNN encoder has a long-term forgetting problem, which cannot effectively capture the global information of the article, and cannot model the "one-to-many" relationship between the article and its keywords well.

Method used

Graph convolutional network (GCN) is used to combine syntactic structure information to convert text into graph structure, learn vector representation of nodes, and guide the decoder to generate diverse keywords through clustering methods to improve the quality of keywords.

Benefits of technology

By introducing structured knowledge, the structural characteristics of the text are extracted, the shortcomings of sequential encoding are supplemented, the model's ability to capture the global characteristics of the text is improved, the higher quality keywords are generated, and the diversity of keywords is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114692605B_ABST
    Figure CN114692605B_ABST
Patent Text Reader

Abstract

The present invention discloses a keyword generation method and device integrating syntactic structure information, which can automatically generate keywords for news articles. The present invention first uses a crawler tool to collect news articles, and constructs a news article data set by manually annotating reference keywords; then preprocesses the text, performs dependency syntactic analysis and filters stop words; then a sequential encoder based on a recurrent neural network and a graph encoder based on a graph convolutional network respectively obtain the contextual semantics and structural features of the article, and uses a clustering method to divide the text into parts containing different sub-topics, and uses multiple decoders based on an attention mechanism to generate keywords in parallel; sampling cross entropy loss is used to optimize model parameters; finally, based on the trained model, automatic keyword generation is performed on the news articles to be processed. The present invention compensates for the long-distance word dependency information loss problem existing in sequential coding through syntactic structure information, thereby improving the quality of generated keywords.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a keyword generation method and device integrating syntactic structure information, belonging to the technical field of Internet and artificial intelligence. Background Art

[0002] With the popularization of computer networks and the development of communication technology, people are exposed to news articles published by various media platforms in social, economic, cultural activities and daily life. In an era of exponential growth in data size, using computer technology to compress information in long news articles and extract the core content can help readers quickly identify valuable information.

[0003] Current keyword generation models are all based on a sequence to sequence (Sequence to Sequence, Seq2Seq) model. This model adopts an encoder-decoder architecture, in which the encoder is responsible for encoding the input text sequence into an intermediate vector, and the decoder decodes from the vector to generate the corresponding input sequence. However, existing methods usually use recurrent neural networks (RNN) to implement encoders and decoders, but RNN has a long-term forgetting problem and cannot effectively capture the global information of the article. On the other hand, the basic Seq2Seq model cannot model the "one-to-many" relationship between the article and its keywords well. To this end, the present invention proposes a keyword generation model based on text syntactic structure information, uses a graph convolutional network (GCN) to mine the deep structural information of the text, and displays the guidance decoder to generate diversified keywords through clustering. Summary of the invention

[0004] In view of the problems and shortcomings in the prior art, the present invention provides a keyword generation method and device integrating syntactic structure information, which converts serialized text into a graph structure through the syntactic structure of the text, and learns the vector representation of the node through GCN, which contains the structural features of the text. The shortcomings of RNN encoding are compensated by structured information, and the clustering method is used to display the guidance model to generate different keywords, thereby improving the quality of keywords.

[0005] To achieve the above-mentioned purpose of the invention, the keyword generation method of the present invention that integrates syntactic structure information first divides the article into sentences and words, and uses a syntactic dependency analysis tool to obtain the syntactic analysis results; then constructs a syntactic graph based on the syntactic analysis structure, maps the text words into nodes in the graph, and the relationship between the words is reflected by the edges; then constructs a sequence and graph encoder to obtain the feature representation of the article; finally, the feature representation is input into the decoder to generate the keywords of the news article. The method of the present invention mainly includes four steps, which are as follows:

[0006] Step 1: News article collection. Use crawler tools to collect news articles from multiple media platforms, accumulate sample data sets, and then filter the sample data sets to reduce sample duplication; manually annotate each sample in the sample set to construct training examples: news articles and standard keywords;

[0007] Step 2: Text preprocessing. The article is divided into sentences and words, and the syntactic analysis results are obtained using the syntactic dependency analysis tool; secondly, a syntactic graph is constructed based on the syntactic analysis structure, and the text words are mapped to nodes in the graph. The relationship between words is reflected through edges;

[0008] Step 3: Train a keyword generation model based on syntactic structure information fusion. First, learn word representations through a dual encoding method of sequential encoding and structural encoding. Then the subgraph clustering network divides the text content according to the meaning of the entire text, thereby constructing a unique sub-topic representation for each decoder. After that, the sequential decoder with attention mechanism generates corresponding keywords based on the generated sub-topic representation; finally, the cross entropy is used as the loss function to optimize the model parameters;

[0009] Step 4: Generate keywords for the news article to be processed. For the news article that needs to predict keywords, first use the syntactic dependency analysis tool to analyze the syntax, then build the text syntax graph, and input the original news article and the syntax graph into the keyword generation model trained in step 3 to generate the keywords of the news article.

[0010] Furthermore, the step 3 includes the following sub-steps:

[0011] Sub-step 3-1, construct the input layer, which receives the text word sequence as input, and uses the pre-trained word2vec model to map each word to the corresponding word vector to obtain the original word vector representation sequence E W ;

[0012] Sub-step 3-2, construct the text encoding layer, using a two-layer BiGRU to encode the word vector sequence E w Perform sequential semantic encoding to obtain the word vector sequence E w The hidden state vector of BiGRU(E w ):

[0013]

[0014]

[0015]

[0016] where u t is word embedding, represents the state vector of the previous GRU unit, Represents the state vector of the next GRU unit;

[0017] The GCN network is used to learn the constructed text graph data; GCN uses neighbor node aggregation to update node information, which is defined as follows:

[0018] H l =ReLU(AH l-1 W l )

[0019] Where A is the adjacency matrix of the text graph, H l Represents the output of the current layer, and initializes each node representation with the word representation, W l is a training parameter; for a graph convolutional network with L layers, the node obtains the information of L-order neighbor nodes, so the feature vector representation of the node has structural information;

[0020] Sub-step 3-3, construct a sub-graph generation layer. Based on the text graph, split and cluster the text graph to obtain multiple sub-graphs containing different aspects of the article; for each node, use the following formula to calculate the probability that the node belongs to each sub-graph:

[0021] assigments = softmax(W a H L +b a )

[0022] Among them, H L represents the output of the last layer of GCN, W a , b a is a learnable parameter, a represents the network for calculating attention weights, and softmax is a normalization function;

[0023] Afterwards, the weighted sum of the node representations can be used to obtain the representation of the subgraph:

[0024]

[0025] Sub-steps 3-4, construct a keyword decoding layer; use multiple identical decoders to decode in parallel to generate keywords; a single decoder is implemented using a unidirectional GRU and combined with a replication mechanism; when decoding time step j, according to the representation u of the previous word j-1 and the previous hidden state s j-1 , calculate the current hidden state:

[0026] s j =GRU(u j-1 ,s j-1 )

[0027] After that, the attention mechanism is used to calculate the attention weight of each word in the input text:

[0028]

[0029] α j =softmax(e j )

[0030] in, represents the feature vector of the ith word in the text sequence calculated by BiGRU, e ij Measures the correlation between the predicted j-th word and the i-th word in the original text, e j Represents the attention weight of the original word when predicting the jth word;

[0031] By weighted summing of word feature vectors, we get the current context representation vector:

[0032]

[0033] Then, combining the subgraph representation, context vector and hidden state, we get the distribution of words in the vocabulary:

[0034] P vocab =softmax(W g [s j ;c j ; g]+b g )

[0035] Among them, g is the network that calculates the vocabulary distribution;

[0036] Finally, at time step j, the final distribution of the predicted word is as follows,

[0037] P final =(1-λ j )·P vocab +λ j ·P copy

[0038] λ j =sigmoid(W λ [c j ;u j-1 ;s j ; g]+b λ )

[0039] Where P copy =α j ,λ j represents the probability of copying a word from the original text, and λ is the network that calculates the copy probability;

[0040] Sub-step 3-5, construct a loss function layer, and use the cross entropy loss between the keywords generated by this layer and the reference keywords as the training loss function of the model; the training loss of this group of samples is obtained according to the following loss function calculation formula:

[0041]

[0042] Where D is the training data set, x is the input text, y is the target keyword, and θ is the parameter set of the model;

[0043] Sub-steps 3-6, training the model; all parameters to be trained are initialized by random initialization, and the Adam optimizer is used to perform gradient back propagation to update the model parameters during the training process. When the training loss no longer decreases or the number of training rounds exceeds a certain number of rounds, the model training ends.

[0044] Furthermore, the syntactic dependency analysis tool is HanLP.

[0045] The present invention also provides a keyword generation device integrating syntactic structure information, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, the keyword generation method integrating syntactic structure information is implemented.

[0046] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0047] (1) The present invention introduces structured knowledge, which can effectively extract the structural features of the text to make up for the shortcomings of sequential coding. Sequential semantics and structural semantics complement each other, effectively improving the model's ability to capture the global feature information of the text, overcoming the problem that the existing keyword generation method cannot effectively capture global information, thereby improving the quality of keyword generation;

[0048] (2) The present invention adopts a clustering method to explicitly divide a news article into multiple parts containing sub-topics, and assigns an independent encoder to each sub-topic. It adopts a multi-decoder parallel decoding method to generate keywords, which can improve the diversity of keywords generated by the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 The figure is a processing flow chart of an embodiment of the present invention.

[0050] Figure 2 The figure is an overall framework diagram of the method according to the embodiment of the present invention. DETAILED DESCRIPTION

[0051] The technical solution provided by the present invention will be described in detail below in conjunction with specific embodiments. It should be understood that the following specific implementation methods are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0052] The present invention provides a keyword generation method integrating syntactic structure information, and its processing flow is as follows: Figure 1 As shown, the implementation framework is as follows Figure 2 As shown, the specific implementation steps include the following:

[0053] Step 1: news article data set collection. Without loss of generality, this embodiment first collects a large number of news articles from the Internet through a crawler tool, and then filters them to reduce the sample duplication rate; and manually annotates training examples: news articles and standard keywords, which together constitute a sample data set D.

[0054] Step 2, data preprocessing, first preprocess each news article in data set D by segmenting words and sentences, then use the syntactic analysis tool HanLP to parse the text to obtain the syntactic dependencies of the text; filter out stop words; construct a syntactic graph based on the syntactic analysis results, map the text words to nodes in the graph, and the relationship between words is reflected by edges.

[0055] Step 3: Use the data set D processed in step 2 to train the keyword generation model integrating syntactic structure information. The implementation of this step can be divided into the following sub-steps:

[0056] Sub-step 3-1, construct the input layer, which receives the text word sequence as input, and uses the pre-trained word2vec model to map each word to the corresponding word vector to obtain the original word vector representation sequence E W .

[0057] Sub-step 3-2, construct the text encoding layer. This embodiment uses a two-layer BiGRU to encode the word vector sequence E w Perform sequential semantic encoding to obtain the word vector sequence E w The hidden state vector of BiGRU(E w ):

[0058]

[0059]

[0060]

[0061] where u t is word embedding, represents the state vector of the previous GRU unit, Represents the state vector of the next GRU unit.

[0062] The GCN network is used to learn and construct text graph data. Through multi-layer GCN, word nodes can not only obtain word information of neighbors, but also obtain word information of more distant words. GCN uses neighbor node aggregation to update node information, which is defined as follows:

[0063] H l =ReLU(AH l-1 W l )

[0064] Where A is the adjacency matrix of the text graph, H l Represents the output of the current layer, and initializes each node representation with the word representation, W l is a training parameter. For a graph convolutional network with L layers, the node obtains the information of L-order neighbor nodes, so the feature vector representation of the node has structural information.

[0065] Sub-step 3-3, construct a sub-graph generation layer. Based on the text graph, split and cluster the text graph to obtain multiple sub-graphs containing different aspects of the article. For each node, use the following formula to calculate the probability that the node belongs to each sub-graph:

[0066] assigments = softmax(W a H L +b a )

[0067] Among them, H L represents the output of the last layer of GCN, W a , b a is a learnable parameter, a represents the network for calculating attention weights, and softmax is a normalization function.

[0068] Afterwards, the weighted sum of the node representations can be used to obtain the representation of the subgraph:

[0069]

[0070] Sub-step 3-4, construct the keyword decoding layer. This embodiment uses multiple identical decoders to decode in parallel to generate keywords. Among them, a single decoder is implemented using a unidirectional GRU and combined with a replication mechanism. At decoding time step j, according to the representation u of the previous word j-1 and the previous hidden state s j-1 , calculate the current hidden state:

[0071] s j =GRU(u j-1 ,s j-1 )

[0072] After that, the attention mechanism is used to calculate the attention weight of each word in the input text:

[0073]

[0074] α j =softmax(e j )

[0075] in, is a learnable parameter, represents the feature vector of the ith word in the text sequence calculated by BiGRU, e ij Measures the correlation between the predicted j-th word and the i-th word in the original text, e j Represents the attention weight (unnormalized) of the original word when predicting the jth word.

[0076] By weighted summing of word feature vectors, we get the current context representation vector:

[0077]

[0078] Among them, H s It is the feature matrix composed of the feature vectors of the original words.

[0079] Then, combining the subgraph representation, context vector and hidden state, we get the distribution of words in the vocabulary:

[0080] P vocab =softmax(W g [s j ;c j ; g]+b g )

[0081] Where g is the network that calculates the word list distribution. In addition to generating words from the word list, copying words from the source text is also a way to generate words. In the copying mechanism, the attention weight of the word can be regarded as the distribution of the generated word in the source text at the current moment.

[0082] Finally, at time step j, the final distribution of the predicted word is as follows,

[0083] P final =(1-λ j )·P vocab +λ j ·P copy

[0084] λ j =sigmoid(W λ [c j ;u j-1 ;s j ; g]+b λ )

[0085] Where P copy =α j ,λ j represents the probability of copying a word from the original text, and g is the subtopic feature vector.

[0086] Sub-step 3-5, construct a loss function layer, and the cross entropy loss between the keywords generated by this layer and the reference keywords is used as the training loss function of the model. The training loss of this group of samples is obtained according to the following loss function calculation formula:

[0087]

[0088] Among them, D is the training data set, x is the input text, y is the target keyword, and θ is the parameter set of the model.

[0089] Sub-step 3-6, training the model. This embodiment uses random initialization to initialize all parameters to be trained. During the training process, the Adam optimizer is used to perform gradient back propagation to update the model parameters, and the initial learning rate is set to 0.001. When the training loss no longer decreases or the number of training rounds exceeds 10 rounds, the model training ends.

[0090] Step 4: Use the trained parameters to initialize the keyword model to generate keywords. The model takes the news article after text preprocessing as input, first uses the syntactic dependency analysis tool HanLP to analyze the syntax, and then constructs the text syntax graph. Specifically, the article is first sequence-encoded and structurally encoded, and then phrases are generated word by word at the decoding layer as target keywords. The initial word is a special start marker " <start>", the predicted word at each moment is the word with the highest probability output by the pointer generation layer, and when the output ends, the marker" <end>", stop generating words and output the generated word sequence as the predicted keywords of the input news article.

[0091] Based on the same inventive concept, an embodiment of the present invention also provides a keyword generation device that integrates syntactic structure information, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, it implements the above-mentioned keyword generation method that integrates syntactic structure information.

[0092] The technical means disclosed in the scheme of the present invention are not limited to the technical means disclosed in the above-mentioned implementation mode, but also include technical schemes composed of any combination of the above-mentioned technical features. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications are also regarded as the protection scope of the present invention.< / end> < / start>

Claims

1. A keyword generation method integrating syntactic structure information, characterized in that: The steps include: Step 1: News article collection Use crawler tools to collect news articles from multiple media platforms, accumulate sample data sets, and then filter the sample data sets to reduce sample duplication; manually annotate each sample in the sample set to construct training examples: news articles and standard keywords; Step 2: Text Preprocessing The article is divided into sentences and words, and the syntactic analysis results are obtained using the syntactic dependency analysis tool; secondly, a syntactic graph is constructed based on the syntactic analysis structure, and the text words are mapped to nodes in the graph. The relationship between words is reflected by edges; Step 3: Train a keyword generation model based on syntactic structure information fusion First, word representations are learned through a dual encoding method of sequential encoding and structural encoding; then the subgraph clustering network divides the text content according to the meaning of the entire text, thereby constructing a unique sub-topic representation for each decoder; then the sequential decoder with attention mechanism generates corresponding keywords based on the generated sub-topic representation; finally, the cross entropy is used as the loss function to optimize the model parameters; including the following sub-steps: Sub-step 3-1, construct the input layer; Sub-step 3-2, constructing a text encoding layer; Sub-step 3-3, construct a sub-graph generation layer. Based on the text graph, split and cluster the text graph to obtain multiple sub-graphs containing different aspects of the article; for each node, use the following formula to calculate the probability that the node belongs to each sub-graph: assigments=softmax(W a H L +b a ) Among them, H L represents the output of the last layer of GCN, W a , b a is a learnable parameter, a represents the network for calculating attention weights, and softmax is a normalization function; Afterwards, the weighted sum of the node representations can be used to obtain the representation of the subgraph: Sub-step 3-4, constructing a keyword decoding layer; Sub-steps 3-5, construct the loss function layer; Sub-steps 3-6, training the model; Step 4: Generate keywords for the news articles to be processed For news articles that need to predict keywords, first use the syntactic dependency analysis tool to analyze the syntax, then build a text syntax graph, and input the original news article and the syntax graph into the keyword generation model trained in step 3 to generate the keywords of the news article.

2. The keyword generation method integrating syntactic structure information according to claim 1, characterized in that: In step 3, Sub-step 3-1 includes the following specific processes: the input layer receives the text word sequence as input, uses the pre-trained word2vec model to map each word to the corresponding word vector, and obtains the original word vector representation sequence E W ; Sub-step 3-2 includes the following specific process: a two-layer BiGRU is used to vectorize the word sequence E w Perform sequential semantic encoding to obtain the word vector sequence E w The hidden state vector of BiGRU(E w ): where u t is word embedding, represents the state vector of the previous GRU unit, Represents the state vector of the next GRU unit; The GCN network is used to learn the constructed text graph data; GCN uses neighbor node aggregation to update node information, which is defined as follows: A l =ReLU(AH l-1 W l ) Where A is the adjacency matrix of the text graph, H l Represents the output of the current layer, and initializes each node representation with the word representation, W l is a training parameter; for a graph convolutional network with L layers, the node obtains the information of L-order neighbor nodes, so the feature vector representation of the node has structural information; Sub-steps 3-4 include the following specific processes: using multiple identical decoders to decode in parallel to generate keywords; wherein a single decoder is implemented using a unidirectional GRU and combined with a replication mechanism; when decoding time step j, according to the representation u of the previous word j-1 and the previous hidden state s j-1 , calculate the current hidden state: with j =GRU(in j-1 ,with j-1 ) After that, the attention mechanism is used to calculate the attention weight of each word in the input text: a j =softmax(e j ) in, represents the feature vector of the ith word in the text sequence calculated by BiGRU, g is the subtopic feature vector, e ij Measures the correlation between the predicted j-th word and the i-th word in the original text, e j Represents the attention weight of the original word when predicting the jth word; By weighted summing of word feature vectors, we get the current context representation vector: Among them, H s The feature matrix is ​​composed of the feature vectors of the original words; Then, combining the subgraph representation, context vector and hidden state, we get the distribution of words in the vocabulary: P vocab =softmax(W g [s j ;c j ;g]+b g ) Among them, g is the network that calculates the vocabulary distribution; Finally, at time step j, the final distribution of predicted words is as follows: P final =(1-λ j )·P vocab +λ j ·P copy λ j =sigmoid(W λ [c j ;u j-1 ;s j ;g]+b λ ) Where P copy =α j ,λ j represents the probability of copying a word from the original text, and λ is the network that calculates the copy probability; Sub-steps 3-5 include the following specific processes: the cross entropy loss between the keywords generated in this layer and the reference keywords is used as the training loss function of the model; the training loss of this group of samples is obtained according to the following loss function calculation formula: Where D is the training data set, x is the input text, y is the target keyword, and θ is the parameter set of the model; Sub-steps 3-6 include the following specific processes: all parameters to be trained are initialized by random initialization, and the Adam optimizer is used to perform gradient backpropagation to update the model parameters during the training process. When the training loss no longer decreases or the number of training rounds exceeds a certain number, the model training ends.

3. The keyword generation method integrating syntactic structure information according to claim 1, characterized in that: The syntactic dependency analysis tool is HanLP.

4. A keyword generation device integrating syntactic structure information, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into a processor, the method for generating keywords integrating syntactic structure information as described in any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Automatic news title generation method

    CN111241816A

  • Conversation problem generation method and system considering emotion and theme, and storage medium

    CN111949761A