An unsupervised keyword extraction method based on gated topic model
Through the unsupervised keyword extraction method based on the gated theme model, adaptively modeling the document semantics and combining the topic importance and correlation design scoring algorithm, the problems of differences in keyword extraction diversity and semantic richness in the existing technology are solved, and more accurate and diversified keyword extraction effects are achieved.
Patent Information
- Application Number
- CN202311341725.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-17
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2043-10-17
AI Technical Summary
Existing unsupervised keyword extraction technology is difficult to effectively capture the semantic richness differences and keyword diversity between documents, resulting in the possible redundancy of extracted keywords and information loss.
The unsupervised keyword extraction method based on the gated theme model is adopted, and the document semantics are adaptively modeled through comparative learning and gating mechanisms, and the keyword scoring algorithm is designed in combination with the topic importance and theme relevance to improve the diversity of keywords.
Semantic adaptive document semantic representation is realized, which improves the diversity and accuracy of keywords, avoids excessive attention to the core topics of the text, and ensures the diversity and information integrity of keywords.
Smart Images

Figure CN117390157B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of artificial intelligence, big data and natural language processing, and specifically relates to an unsupervised keyword extraction method based on a gated topic model. Background Art
[0002] Faced with the rapid generation and widespread dissemination of electronic text data, how to accurately and efficiently extract key information from massive text data has become an urgent need. Keyword extraction technology can not only help people quickly understand the core points of text content, but also provide support for downstream applications such as information retrieval, text summarization, and topic modeling, which has important research significance.
[0003] Keyword extraction technology can be divided into supervised methods and unsupervised methods. Since annotated keywords are usually difficult to obtain in real scenarios, unsupervised keyword extraction is often more practical in applications. Existing unsupervised keyword extraction technologies can generally be classified into three paradigms: Statistical feature-based keyword extraction calculates keyword scores by defining and selecting some statistical features, such as word frequency and word position. Although the implementation is relatively simple, this type of method essentially treats words as symbols and attempts to find the connection between keywords through statistical features without modeling keywords from a semantic level. Therefore, there are certain limitations in capturing text semantics and topics. Graph-based keyword extraction methods convert original documents into graph structures, in which the relationships between words constitute the edges of the graph. Keywords are screened from complex networks by solving graph optimization problems. This method can capture the local context information of words, but it also does not model the semantics of text and words, and cannot effectively capture global semantic information. Embedding-based keyword extraction is the most advanced method in the field of unsupervised keyword extraction. This method uses a pre-trained language model to encode documents and candidate words, and then ranks the candidate words according to the semantic similarity score between the document embedding and the candidate word embedding. This semantic-based method significantly improves the accuracy of keyword extraction, but it also has some problems. First, it is difficult for this method to capture the differences in semantic richness between documents because it represents document semantics as fixed embeddings. Second, using only semantic similarity metric scoring may cause keywords to be too concentrated in one topic, resulting in insufficient diversity. Summary of the invention
[0004] 1. Technical issues to be resolved
[0005] The technical problem to be solved by the present invention is how to provide an unsupervised keyword extraction method based on a gated topic model to solve two problems: first, there are wide differences in the semantic richness between documents, and an adaptive document semantic modeling method is required; second, the keywords of a single document are usually distributed under multiple topics, and only focusing on core topics will lead to keyword redundancy and information loss, and it is necessary to improve the diversity of keyword extraction.
[0006] (II) Technical solution
[0007] In order to solve the above technical problems, the present invention proposes an unsupervised keyword extraction method based on a gated topic model, which comprises the following steps:
[0008] Step 1: Word segmentation and part-of-speech tagging
[0009] Before encoding the input text, the original natural language text data needs to be preprocessed;
[0010] Step 2: Noun phrase extraction
[0011] Based on the POS tagging results, only noun phrases in the original text were retained as candidate keywords;
[0012] Step 3: Document encoding and candidate word representation
[0013] Encode document words and candidate keywords based on GloVe embedding to obtain word embedding representation;
[0014] Step 4: Topic Modeling
[0015] S41. First, for any document d in the corpus, use the word embedding obtained in step 3 to construct the context vector representation z of d d ;
[0016] S42. From the perspective of topic modeling, the document is represented as the weighted sum of topic embeddings, and then the document context is represented as z d Refactoring to another representation of the subject
[0017] S43, after obtaining the document context vector representation z d and its topic representation d Afterwards, the contrastive learning strategy is used to optimize the model parameters. The goal of contrastive learning is to minimize the loss function
[0018] S44, to minimize To train the topic model for the goal, a set of topic representations M is extracted from the entire corpus T ={m 1 ,m2 ,…,m K}, and determine the weight vector p of each input document on these K topics d ={w 1 ,w 2 ,…,w K};
[0019] Step 5: Keyword extraction
[0020] For each candidate word np i , calculate its score on K topics, np i The final score of is the maximum value of these K scores. All candidate words are sorted according to the final score, and the top N candidate words are extracted as the keywords of document d.
[0021] (III) Beneficial effects
[0022] The present invention proposes an unsupervised keyword extraction method based on a gated topic model. The present invention discloses an unsupervised keyword extraction method based on a gated topic model, and its main advantages are embodied in the following aspects:
[0023] (1) We propose a semantically adaptive document semantic representation method that trains a neural topic model on the entire corpus to mine relevant topics in the field, and adopts a gating mechanism to independently weight document topics so that documents with higher semantic richness are assigned relatively more topics.
[0024] (2) A new keyword scoring algorithm is designed using document topic information, which takes into account the impact of topic similarity and topic importance on keyword evaluation. By compromising these two factors, excessive attention to the core topic of the text is avoided, thereby improving the diversity of the extracted keywords. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is the overall framework diagram of the method of the present invention;
[0026] Figure 2 This is a framework diagram of the gated topic model of the present invention. DETAILED DESCRIPTION
[0027] In order to make the purpose, content and advantages of the present invention more clear, the specific implementation methods of the present invention are further described in detail below in conjunction with the drawings and examples.
[0028] This paper proposes an unsupervised keyword extraction method suitable for complex document input in real scenarios. It models the semantic richness of documents through a set of adaptively assigned document topics, and then extracts more diverse keywords based on topic information. It achieves the extraction of more diverse keyword information based on accurate understanding of the semantics of the source document.
[0029] The present invention provides an unsupervised keyword extraction method based on a gated topic model, the method comprising the following steps:
[0030] Step 1: Word segmentation and part-of-speech tagging
[0031] Before encoding the input text, the original natural language text data needs to be preprocessed;
[0032] Step 2: Noun phrase extraction
[0033] Based on the POS tagging results, only noun phrases in the original text were retained as candidate keywords;
[0034] Step 3: Document encoding and candidate word representation
[0035] Encode document words and candidate keywords based on GloVe embedding to obtain word embedding representation;
[0036] Step 4: Topic Modeling
[0037] S41. First, for any document d in the corpus, use the word embedding obtained in step 3 to construct the context vector representation z of d d ;
[0038] S42. From the perspective of topic modeling, the document is represented as the weighted sum of topic embeddings, and then the document context is represented as z d Refactoring to another representation of the subject
[0039] S43, after obtaining the document context vector representation z d and its topic representation d Afterwards, the contrastive learning strategy is used to optimize the model parameters. The goal of contrastive learning is to minimize the loss function
[0040] S44, to minimize To train the topic model for the goal, a set of topic representations M is extracted from the entire corpus T ={m 1 ,m 2 ,…,m K}, and determine the weight vector p of each input document on these K topics d ={w1 ,w 2 ,…,w K};
[0041] Step 5: Keyword extraction
[0042] For each candidate word np i , calculate its score on K topics, np i The final score of is the maximum value of these K scores. All candidate words are sorted according to the final score, and the top N candidate words are extracted as the keywords of document d.
[0043] like Figure 1 The figure shows the overall model framework diagram. The unsupervised keyword extraction method based on the gated topic model includes five steps: word segmentation and part-of-speech tagging, noun phrase extraction, document encoding and candidate word representation, topic modeling, and keyword extraction.
[0044] Step 1: Word segmentation and part-of-speech tagging
[0045] Before encoding the input text, the original natural language text data needs to be preprocessed. Given any document d in the data set, first perform a text segmentation operation on it to obtain a word sequence {t 1 ,t 2 ,...,t n Then, the word sequence after word segmentation is annotated with part-of-speech (POS), and each word is assigned an appropriate POS tag.
[0046] Step 2: Noun phrase extraction
[0047] Based on the POS tagging results, only noun phrases in the original text are retained as candidate keywords in this step. Generally speaking, nouns are usually used as the main body of sentences such as subjects and objects, and the expression is more semantically rich, while adjectives and adverbs are mostly used as modifiers to express a degree or polarity, with low information content and are not suitable to appear as keywords alone. Specifically, the present invention extracts phrases that meet the regular expression {<NN.*|JJ> *<NN.*>} pattern as candidate keywords, where NN represents a noun part-of-speech tag, JJ represents an adjective part-of-speech tag, and the overall meaning is a combination of zero or more adjectives and at least one noun. The resulting candidate keyword set is defined as C = {np 1 ,np 2 ,...,np m}.
[0048] Step 3: Document encoding and candidate word representation
[0049] The present invention encodes document words and candidate keywords based on GloVe embedding to obtain word embedding representation.
[0050] For document words, the specific encoding process can be formally expressed as a formula:
[0051] {e 1 ,e 2 ,…,e n}=GloVe({t 1 ,t 2 ,…,t n}) (1)
[0052] in It is the word t i The embedding representation of , v represents the dimension of word embedding.
[0053] For a candidate keyword, the present invention directly calculates the average value of the embedded representations of the words it contains as its representation:
[0054]
[0055] in, It is the candidate keyword np j The embedded representation of .
[0056] Step 4: Topic Modeling
[0057] The present invention formalizes the topic modeling process as an unsupervised problem similar to autoencoding, and adopts a contrastive learning strategy to train the model. In order to enable the topic model to adaptively model the semantic richness of the text, the present invention further represents the document semantics as a set of weights for each topic, and introduces a gating mechanism to independently regulate the weight of each topic. The overall framework of the topic model is as follows: Figure 2 shown.
[0058] S41. First, for any document d in the corpus, use the word embedding obtained in step 3 to construct the context vector representation z of d d , the calculation process is as follows:
[0059]
[0060] in, Represents the contextual embedding representation of document d, represented by word embedding e i The weighted sum of , i = 1, 2, ..., n represents the position index of the word in the document.
[0061] In order to make the vector representation capture the core semantic information of the document, the present invention uses TF-IDF weights instead of attention weights as the summation coefficient of word embedding. This is mainly due to two factors. On the one hand, TF-IDF is simpler and more general. On the other hand, TF-IDF weights imply the information of the entire corpus, which can alleviate the domain-independent common vocabulary pairs to a certain extent. d To cause adverse effects. Specifically, the word t i The TF-IDF weight is expressed as tfidf(t i ), the calculation process is as follows:
[0062]
[0063] Among them, n i and n represent the word t respectively i The number of occurrences in document d and the entire corpus. i ∈d j}| indicates that the entire corpus contains word t i N represents the total number of documents in the corpus. To avoid division by zero, the denominator of the formula is adjusted to 1+|{j:t i ∈d j}|.
[0064] S42. From the perspective of topic modeling, a document can also be represented as a weighted sum of topic embeddings, and then the document context is represented as z d Refactoring to another representation of the subject Such as the formula:
[0065] r d =M·p d (5)
[0066] in, M is a randomly initialized topic embedding matrix whose parameters are updated during the training phase. T Each row in is defined as a specific topic representation, whose dimension is consistent with the word embedding. K is a manually set hyperparameter representing the number of topics. d =(w 1 ,w 2 ,…,w K ) represents the weight vector of document d on K topics, and its specific value is determined by formula (6):
[0067] p d =sigmoid(W·z d +b) (6)
[0068] Where W and b are two learnable parameters.
[0069] Different topics should not be in a completely competitive relationship, that is, a document with rich semantics can contain multiple topics with relatively high weights at the same time. Relatively speaking, a document with low semantic richness should have low weights for all topics. Therefore, the present invention uses a gating mechanism as shown in formula (6) to assign topics to document d, replacing the original softmax(·) with sigmoid(·) to cancel the competitive constraint relationship between document topics, so that its topic weight vector p d Each component in can be adaptively determined according to the semantic richness of the document.
[0070] S43, after obtaining the document context vector representation z d and its topic representation d Afterwards, the contrastive learning strategy is used to optimize the model parameters. The goal of contrastive learning is to minimize the loss function For each input document d in the document library D, r d and z d In essence, it is a characterization of d from different perspectives. The two representations should be as similar as possible. For other documents in the corpus that are not related to d, let their context be represented as c, then c should be similar to r. d Therefore, the loss function shown in formula (7) Expected to reduce r d and z d The distance between them, while maximizing r d and Documentation The distance between i , are irrelevant documents randomly sampled from a batch of inputs, defined as negative samples of d.
[0071] In addition, in order to ensure the diversity of topics and avoid redundant topics, the present invention additionally defines a regularization loss to constrain the relationship between topics. The specific definition is shown in formula (8), where is the normalized topic embedding matrix, each column vector of which is normalized to a unit vector, is the identity matrix.
[0072] Finally, the present invention defines the loss function as and The sum of is shown in formula (9), where λ ort is a hyperparameter that controls the weight of the regularization term.
[0073]
[0074]
[0075]
[0076] S44, to minimize The topic model is trained for the purpose of training. The topic modeling method proposed in this invention can extract a set of topic representations M from the entire corpus. T ={m 1 ,m 2 ,…,m K}, and determine the weight vector p of each input document on these K topics d ={w 1 ,w 2 ,…,w K}.
[0077] Step 5: Keyword extraction
[0078] The present invention believes that the importance of keywords is related to two factors: 1) topic importance, that is, the importance of the topic to which the keyword belongs, which can be represented by the weight of the document on the topic in the present invention; 2) topic relevance, that is, the relevance between the keyword and the topic, which can be determined by calculating the similarity between topic embedding and word embedding. In order to integrate the two factors of topic importance and topic relevance to extract diversified keywords, the present invention designs the keyword score function as shown in formula (10):
[0079]
[0080] in, is the indicator function for topic selection, when w j >α and When the value is 1, it means that the subject j is selected, otherwise it is equal to 0. j represents the weight of document d about topic j or the importance of topic j in d. The topic relevance is defined as the candidate word embedding and the theme embedding j The dot product of α, β, λ s These are three adjustable hyperparameters that control the topic importance threshold, topic relevance threshold, and score item sum weight.
[0081] For each candidate word np i The present invention calculates the scores of K topics by the above method. i ,m j ),1≤j≤K. np i The final score of is the maximum value of these K scores, as shown in formula (11):
[0082]
[0083] Finally, according to the final score All candidate words are sorted, and the top N candidate words are extracted as keywords of document d. The complete keyword extraction process is shown in Table 1. Under this design, whether a candidate word can be identified as a keyword will be affected by both the topic importance and the topic relevance, which ensures that candidate words belonging to secondary topics but with high topic relevance can also receive high attention, thereby ensuring the diversity of keywords.
[0084] Table 1 Keyword extraction algorithm integrating topic importance and topic relevance
[0085]
[0086] The present invention discloses an unsupervised keyword extraction method based on a gated topic model, the main advantages of which are embodied in the following aspects:
[0087] (1) We propose a semantically adaptive document semantic representation method that trains a neural topic model on the entire corpus to mine relevant topics in the field, and adopts a gating mechanism to independently weight document topics so that documents with higher semantic richness are assigned relatively more topics.
[0088] (2) A new keyword scoring algorithm is designed using document topic information, which takes into account the impact of topic similarity and topic importance on keyword evaluation. By compromising these two factors, excessive attention to the core topic of the text is avoided, thereby improving the diversity of the extracted keywords.
[0089] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. An unsupervised keyword extraction method based on a gated topic model, characterized in that: The method comprises the following steps: Step 1: Word segmentation and part-of-speech tagging Before encoding the input text, the original natural language text data needs to be preprocessed; Step 2: Noun phrase extraction Based on the POS tagging results, only noun phrases in the original text were retained as candidate keywords; Step 3: Document encoding and candidate word representation Encode document words and candidate keywords based on GloVe embedding to obtain word embedding representation; Step 4: Topic Modeling S41. First, for any document d in the corpus, use the word embedding obtained in step 3 to construct the context vector representation z of d d ; S42. From the perspective of topic modeling, a gating mechanism is used to assign the topic of document d. The document is represented as the weighted sum of topic embeddings, and then the document context is represented as z. d Refactoring to another representation of the subject S43, after obtaining the document context vector representation z d and its topic representation d Afterwards, the contrastive learning strategy is used to optimize the model parameters. The goal of contrastive learning is to minimize the loss function S44, to minimize To train the topic model for the goal, a set of topic representations M is extracted from the entire corpus T ={m1,m2,…,m K }, and determine the weight vector p of each input document on these K topics d ={w1,w2,…,w k }; Step 5: Keyword extraction Considering the influence of topic similarity and topic importance on keyword evaluation, for each candidate word np i , calculate its score on K topics, np i The final score of is the maximum value of these K scores. All candidate words are sorted according to the final score, and the top N candidate words are extracted as the keywords of document d.
2. The unsupervised keyword extraction method based on the gated topic model as claimed in claim 1, characterized in that: The step 1 specifically includes: given any document d in the data set, firstly perform a text segmentation operation on it to obtain a word sequence {t1, t2, ..., t n }; Then, the word sequence after segmentation is annotated with part-of-speech (POS) and each word is assigned an appropriate POS tag.
3. The unsupervised keyword extraction method based on the gated topic model as claimed in claim 2, characterized in that: The step 2 specifically includes: extracting the regular expression {<NN.*|JJ> *<NN.*> All phrases of the pattern are taken as candidate keywords, where NN represents a noun part-of-speech tag, JJ represents an adjective part-of-speech tag, and the overall meaning is a combination of zero or more adjectives and at least one noun; the resulting candidate keyword set is defined as C = {np1, np2, ..., np m }.
4. The unsupervised keyword extraction method based on the gated topic model as claimed in claim 3, characterized in that: In step 3, for the document words, the specific encoding process is expressed as a formula: {e1,e2,…,e n ]=GloVe({t1,t2,…,t n }) (1) in It is the word t i The embedding representation of , v represents the dimension of word embedding.
5. The unsupervised keyword extraction method based on the gated topic model as claimed in claim 4, characterized in that: In step 3, for a candidate keyword, the average value of the embedded representations of the words contained in it is calculated as its representation: in, It is the candidate keyword np j The embedded representation of .
6. The unsupervised keyword extraction method based on the gated topic model as claimed in claim 5, characterized in that: The S41 specifically includes: in, Represents the contextual embedding representation of document d, represented by word embedding e i The weighted sum of , i = 1, 2, ..., n represents the position index of the word in the document; Word t i The TF-IDF weight is expressed as tfidf(t i ), the calculation process is as follows: Among them, n i and n represent the word t respectively i The number of occurrences in document d and the entire corpus; |{j:t i ∈d j }| indicates that the entire corpus contains word t i , N represents the total number of documents in the corpus; to avoid division by zero errors, the denominator of the formula is adjusted to 1+|{j:t i ∈d j }|.
7. The unsupervised keyword extraction method based on the gated topic model as claimed in claim 6, characterized in that: The S42 includes: r d =M·p d (5) in, is a randomly initialized topic embedding matrix whose parameters are updated during the training phase; M T Each row in is defined as a specific topic representation, whose dimension is consistent with the word embedding; K is a manually set hyperparameter representing the number of topics; p d =(w1,w2,…,w K ) represents the weight vector of document d on K topics, and its specific value is determined by formula (6): p d =sigmoid(W·z d +b) (6) Where W and b are two learnable parameters.
8. The unsupervised keyword extraction method based on the gated topic model as claimed in claim 7, characterized in that: The S43 specifically includes: for each input document d in the document library D, r d and z d In essence, it is a characterization of d from different perspectives. The two representations should be as similar as possible. For other documents in the corpus that are not related to d, let their context be represented as c, then c should be similar to r. d As different as possible; therefore, the loss function shown in formula (7) Expected to reduce r d and z d The distance between them, while maximizing r d and documents c1,c2,..., The distance between i , is an irrelevant document randomly sampled from a batch of inputs, defined as a negative sample of d; In addition, in order to ensure the diversity of topics and avoid redundant topics, a regularization loss is defined to constrain the relationship between topics; the specific definition is shown in formula (8), where is the normalized topic embedding matrix, each column vector of which is normalized to a unit vector, is the identity matrix; Finally, the loss function is defined as and The sum of is shown in formula (9), where λ ort is a hyperparameter that controls the weight of the regularization term; 9. The unsupervised keyword extraction method based on the gated topic model as claimed in claim 8, characterized in that: For each candidate word np i , the scores of K topics are calculated as follows: The keyword score function is designed as shown in formula (10): in, is the indicator function for topic selection, when w j >α and When the value is 1, it means that the subject j is selected, otherwise it is equal to 0; w j represents the weight of document d about topic j or the importance of topic j in d; topic relevance is defined as the candidate word embedding and theme embedding j The dot product of α, β, λ s These are three adjustable hyperparameters that control the topic importance threshold, topic relevance threshold, and score item sum weight.
10. The unsupervised keyword extraction method based on the gated topic model according to claim 9, characterized in that: The np i The final score is the maximum value of these K scores, including: For each candidate word np i , calculate the score Score(np) for K topics i ,m j ),1≤j≤K;np i The final score of is the maximum value of these K scores, as shown in formula (11):
Citation Information
Patent Citations
Key phrase generation method and device based on pre-training model and storage medium
CN113934837A
Method and device for extracting keyword represented by topic constraint
CN115687576A