Key Phrase Method and System for Extracting Combined Global and Local Information in Document Hierarchy
Through the document hierarchy combined with global and local information extraction method, the embedding processing and similarity calculation are used to use the BERT and SimCSE models to solve the problem of low accuracy of key phrase extraction in the prior art, and a higher quality key phrase extraction is achieved.
Patent Information
- Application Number
- CN202210697632.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-20
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-06-20
AI Technical Summary
In the prior art, there are problems such as semantic loss, preference for long phrases, and insufficient mining of subject information, resulting in low accuracy in key phrase extraction.
The document hierarchy combined with global and local information extraction method is used to perform word segmentation and part-of-speech annotation through the StandfordCoreNLP tool to generate a collection of candidate key phrases; the BERT and SimCSE models are used for embedding processing, the global similarity and local similarity are calculated, and the topic division and clustering are performed through topic centering, and the key phrases are finally obtained through comprehensive evaluation and sorting.
It significantly improves the accuracy of key phrase extraction, especially when processing long texts, can better retain semantic information, reduces preference for long phrases, and improves semantic diversity of candidate key phrases.
Smart Images

Figure CN115017903B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text analysis, and particularly to a method and system for extracting key phrases by combining global and local information in a document hierarchy. Background Art
[0002] Key phrases are phrases that provide a concise summary of the core content in a document, which can help readers understand the content of the article in a short time. Due to their concise and accurate expression, key phrases are widely used in information retrieval, document classification, recommendation, and search. Embedding-based methods are widely used in unsupervised key phrase extraction tasks. Usually, these methods simply calculate the similarity between phrase embeddings and document embeddings, and there is room for improvement in terms of the practicality and effectiveness of the methods.
[0003] A large number of scholars have conducted research on keyword extraction from text. The related methods can generally be divided into unsupervised extraction methods and supervised extraction methods.
[0004] Supervised methods [Sterckx et al., 2016; Alzaidy et al., 2019; Sun et al., 2020; Mu et al., 2020] usually regard key phrase extraction as a binary classification problem. They not only require a large amount of annotated training data but also always perform poorly when transferred to different domains or types of data sets. Compared with supervised methods, unsupervised methods extract phrases based on the information in the input document itself, which is more general and adaptable. Therefore, in this patent, we focus on unsupervised key phrase extraction models.
[0005] Unsupervised key phrase extraction has been studied by a large number of scholars. Recently, with the progress of text representation learning, embedding-based methods such as EmbedRank [Bennani-Smires 2018] and SIFRank [Sun 2020] have achieved good results. Usually, these methods embed candidate phrases and texts through static pre-trained models like Word2Vec or dynamic pre-trained models like BERT, then calculate the embedding similarity between candidate phrases and the entire text, and sort them according to the scores. Although embedding-based methods perform better than traditional methods based on statistics (e.g., TF-IDF [Salton G, 1975]) and graph-based methods (e.g., PositionRank [Corina Florescu 2017]), simply calculating the similarity between candidate phrases and the full text cannot capture different types of contexts. CCRank [Liang et al., 2021] first proposed jointly modeling global information and local information for keyword extraction, but there are still two problems with his method. First, due to the limitation of the BERT model, his method will automatically truncate the first 512 tokens on long texts, which will lead to a large amount of semantic loss. Second, since the full text vector and the candidate phrase vector are not aligned in the semantic space, his global similarity will give higher scores to phrases with longer candidate phrases, which will cause the model to prefer longer candidate phrases more. Third, he simply uses boundary features to model local information and does not fully explore the topic information of the article. After embedding the key phrases and the full text vector, the present invention visualizes and displays them as follows Figure 4As shown in the figure. The five-pointed star is the vector embedding of the article. Nodes with bold and close colors belong to the same theme. Previous embedding-based methods only considered global similarity, that is, they only selected candidate phrases inside the black dashed box. This obviously did not consider the importance of the local theme of the article. However, phrases from the boundary usually only represent a small part of the article's theme and cannot fully explore the theme information between candidate key phrases. In addition, previous phrases did not consider the limitation that BERT can only encode 512 tokens. Therefore, when facing long texts, the truncation method is usually adopted, and only candidate key phrases of the title + abstract are obtained, without obtaining the key phrases in the conclusion part, resulting in insufficient semantic information and poor performance. The existing invention patent document "A Method, Apparatus, Computer Device, and Storage Medium for Keyword Extraction" with the publication number CN111160017A inputs the text data to be processed into a keyword extraction network model trained by using sequence annotation samples carrying set encoding, and can fully explore the semantic relevance of the context through standard keywords, improving the accuracy of keyword extraction. This application also provides a method, apparatus, computer device, and storage medium for speech evaluation. By inputting the speech to be evaluated into the trained keyword extraction network model, it is possible to extract keywords in speeches that are only relevant to the business for different business scenarios. The specification of this existing patent document also discloses that an initial keyword extraction network model composed of ERNIE-BiLSTM-CRF three-layer network units is deployed in the server 104. The ERNIE network unit is an improved version based on the BERT model, which is optimized for Chinese word-level tasks and has better effects on Chinese entity and entity relationship extraction. The main structure of the model is the same as that of the BERT model and is composed of 12 encoder layers. This existing patent document does not fully disclose the technical solution of this application and cannot achieve the technical effects of this application. The method disclosed in the existing invention patent document "A Method, Apparatus, and Storage Medium for Theme Extraction Oriented to Scientific and Technological Needs" with the publication number CN113255340A includes: obtaining scientific and technological demand text data carrying the first-level theme category label of the industry field; respectively obtaining word vectors and document vectors based on the scientific and technological demand text data belonging to the same first-level theme category; using a deep learning-based theme model to obtain theme word vector representations and theme word sets based on the word vectors and document vectors; clustering the scientific and technological demand text data based on the theme word vectors according to a predetermined number of clusters; using a text sorting algorithm to extract the theme words in the theme word set as keywords and sort the extracted theme words, screening out the theme words as the second-level clustering theme category label words according to the theme word scores, and taking the theme word with the highest score as the second-level theme representative of this category. This existing patent document does not fully disclose the technical solution of this application and cannot achieve the technical effects of this application.
[0006] In summary, the prior art has technical problems such as semantic loss, preference for long and short phrases, and insufficient mining of subject information, resulting in low accuracy of key phrase extraction. Summary of the Invention
[0007] The technical problem to be solved by the present invention is how to solve the technical problems in the prior art, including semantic loss, preference for long and short phrases, and insufficient mining of subject information, resulting in low accuracy of key phrase extraction.
[0008] The present invention solves the above technical problems by adopting the following technical solutions: A method for extracting key phrases by combining document hierarchy and global-local information includes:
[0009] S1. Use the StanfordCoreNLP tool to perform word segmentation and part-of-speech tagging on the input document, and perform NP chunking according to preset extraction rules to generate a set of candidate key phrases;
[0010] S2. Determine whether the length of the input document is less than or equal to a preset document length threshold. If so, use the BERT model to embed and process the input document to obtain a vector representation. If not, obtain the specified range content of the input document according to a preset range, and input the specified range content into the SimCSE model to perform embedding to obtain the vector representations of the candidate key phrases, the title vector, and the ending vector;
[0011] S3. Process the title vector and the ending vector to perform global similarity measurement on the candidate key phrases, and obtain the global similarity accordingly;
[0012] S4. Use the topic centrality to perform topic partitioning and clustering on the candidate key phrases in the full text of the input document according to a preset logic, and obtain the local similarity according to the local similarity evaluation. Among them, the step S4 further includes:
[0013] S41. Use the candidate key phrases as nodes and the similarity between the nodes as edges to construct a complete undirected graph;
[0014] S42. Set an adaptive noise filtering threshold according to the maximum and minimum values of each input document;
[0015] S43. Update the weights of the edges according to the adaptive noise filtering threshold to obtain local significance data, and obtain updated edges according to the local significance data;
[0016] S44. Obtain the position information of the candidate key phrases in the full text of the input document;
[0017] S45. Calculate the local similarity according to the position information;
[0018] S5. Combine and process the location information, the global similarity, and the local similarity to comprehensively evaluate and score the candidate key phrases, and sort the candidate key phrases accordingly to obtain key phrase ranking data;
[0019] S6. Obtain a candidate key phrase sorting data set according to the key phrase ranking data, perform post-processing operations on the candidate key phrases, delete subsets of the candidate key phrase set to obtain semantically diverse key phrases, obtain lexical frequency data, and accordingly remove high-frequency general phrases from the candidate key phrase sorting data set to filter out interference from high-frequency invalid phrases.
[0020] The method proposed by the present invention has greatly improved key phrase extraction. It can be seen from the comparison experiment results that both the global similarity and the local similarity proposed by the present invention have played a role. In the calculation of the local similarity, the noise filtering threshold θ proposed by the present invention. In addition, the model proposed by the present invention has made great progress on long texts, which benefits from the fact that the present invention makes full use of the hierarchical structure of the document. The present invention makes the candidate phrases have higher semantic diversity through diversity operations, making the results more acceptable.
[0021] In a more specific technical solution, step S2 includes:
[0022] S21. Insert a CLS marker at the start position of the input document and a SEP marker at the end position using the BERT model;
[0023] S22. Embed and learn the input document to obtain the vector of each token:
[0024] {H 1 ,H 2 ,…,H n} = BERT({T 1 ,T 2 ,…,T n})
[0025] S23. Then obtain the vector representation of the candidate key phrases according to the preset extraction rules to obtain the candidate phrase vector set:
[0026]
[0027] S24. Feed the title and the end of the input document into the BERT model to obtain the title vector H title and the end vector H end .
[0028] S25. Input the conclusion part and the abstract part of the input document into the BERT model respectively for embedding operations to obtain the vector representation;
[0029] S26 Use the SimCSE model to perform expressions on the input document for long texts.
[0030] In view of the problem that the encoding length of BERT is limited, resulting in that traditional embedding-based methods can only truncate long texts, leading to a large amount of semantic loss, the present invention proposes, according to the document hierarchy and human writing habits, to group the title and abstract as one group and the conclusion as another group, and send them into the pre-trained model for embedding in two times when facing long texts. In this way, both time and space are saved, and the semantic information of the full text can be maximally preserved.
[0031] In view of the deficiencies of traditional methods in processing long texts, the present invention creatively proposes to use SimCSE to segment and encode the abstract and conclusion of the document, so that the model proposed by the present invention can fully learn the information of the document and obtain higher-quality key phrases.
[0032] In a more specific technical solution, in step S3, the title vector H title and the ending vector H end are processed according to the following logic to obtain the global similarity of each candidate key phrase i:
[0033]
[0034] where, ‖.‖ represents the Manhattan distance, represents the global similarity between the candidate phrase i and the entire document.
[0035] In view of the problem of preference for long and short phrases caused by semantic space alignment, according to human writing habits, the title and the last sentence of the ending are used to replace the traditional full-text vector, so as to solve the problem of high scores for long and short phrases.
[0036] In a more specific technical solution, step S42 includes:
[0037] S421. Use the graph centrality calculation method to process the candidate key phrase i according to the following logic:
[0038]
[0039] where,
[0040] S422. Set the self-adaptive noise filtering threshold θ according to the following logic;
[0041] θ = min(e ij ) + β × (max(e ij ) - min(e ij ))
[0042] The present invention performs a post - processing operation on candidate phrases. A threshold is set to filter out each candidate phrase that appears in the top 20% frequently in a specific field, avoiding the interference of high - frequency invalid words, and then the semantic diversity of candidate key phrases is improved by deleting subsets.
[0043] In a more specific technical solution, step S43 includes:
[0044] S431. Obtain the local saliency data by using the following logic:
[0045]
[0046] wherein, represents the local saliency of candidate phrase i;
[0047] S432. Obtain the updated edge according to the local saliency data, and when the weight of the updated edge is less than 0, set the weight of the updated edge to 0.
[0048] In a more specific technical solution, step S44 includes:
[0049] S441. Calculate the position where the candidate key phrase first appears in the input document by using the following logic as the position score of the candidate key phrase:
[0050]
[0051] where p 1 is the position where candidate term i first appears;
[0052] S442. Smooth - process the position score of the candidate key phrase by using the softmax function to obtain the position information by using the following logic:
[0053]
[0054] In a more specific technical solution, in step S45, the position information is processed by using the following logic to obtain the local similarity of the candidate key phrase i
[0055]
[0056] In the local text information modeling of the present invention, the theme centrality is adopted, so that the theme information in the full text can be identified, and the local theme information can be captured better compared with the boundary centrality.
[0057] In a more specific technical solution, step S5 includes:
[0058] S51. Use the following logic to perform a multiplicative synthesis process on the global similarity and the local similarity of the candidate key phrases, and thereby obtain a candidate key phrase score:
[0059]
[0060] S52. Sort the candidate key phrases according to the candidate key phrase scores to obtain the key phrase ranking data.
[0061] In a more specific technical solution, in step S6, according to the fine-grained key phrases, the coarse-grained key phrases are deleted, and thereby the semantic diversity key phrases are obtained.
[0062] In a more specific technical solution, the document hierarchy combined global-local information extraction key phrase system includes:
[0063] A candidate phrase generation module that uses the StandfordCoreNLP tool to perform word segmentation and part-of-speech tagging on the input document, and performs NP chunking according to preset extraction rules to generate a candidate key phrase set;
[0064] A BERT model embedding module for determining whether the length of the input document is less than or equal to a preset document length threshold. If so, use the BERT model to perform embedding processing on the input document to obtain a vector representation. If not, obtain the specified range content of the input document according to a preset range, and input the specified range content into the SimCSE model to perform embedding to obtain the vector representations of the candidate key phrases, the title vector, and the ending vector. The SimCSE model embedding module is connected to the candidate phrase generation module;
[0065] A global similarity measurement module for processing the title vector and the ending vector to perform global similarity measurement on the candidate key phrases, and thereby obtain a global similarity. The global similarity measurement module is connected to the BERT model embedding module;
[0066] A local similarity evaluation module for using the topic centrality to perform topic division and clustering on the candidate key phrases in the entire input document according to preset logic, and thereby obtaining a local similarity through local similarity evaluation. The local similarity evaluation module is connected to the candidate phrase generation module. Among them, the local similarity evaluation module further includes:
[0067] An undirected graph construction module for using the candidate key phrases as nodes and the similarity between the nodes as edges to construct a complete undirected graph;
[0068] A noise filtering threshold setting module for setting an adaptive noise filtering threshold according to the maximum value and the minimum value of each input document;
[0069] A noise filtering module, configured to update the weight of the edge according to the adaptive noise filtering threshold update to obtain local saliency data, and obtain an updated edge according to the local saliency data. The noise filtering module is connected to the undirected graph construction module and the noise filtering threshold setting module;
[0070] A position acquisition module, configured to obtain the position information of the candidate key phrases in the full text of the input document according to the new complete undirected graph;
[0071] A local similarity calculation module, configured to calculate the local similarity according to the position information. The local similarity calculation module is connected to the position acquisition module;
[0072] A key phrase ranking module, configured to comprehensively process the position information, the global similarity, and the local similarity to comprehensively evaluate and score the candidate key phrases, and sort the candidate key phrases accordingly to obtain key phrase ranking data. The key phrase ranking module is connected to the global similarity metric module and the local similarity evaluation module;
[0073] A post-processing module, configured to obtain a candidate key phrase sorting data set according to the key phrase ranking data, perform post-processing operations on the candidate key phrases, delete a subset of the candidate key phrase set to obtain semantically diverse key phrases, obtain word frequency data, and accordingly remove high-frequency common phrases from the candidate key phrase sorting data set to filter out interference from high-frequency invalid phrases. The post-processing module is connected to the key phrase ranking module.
[0074] The present invention has the following advantages compared with the prior art: The method proposed by the present invention greatly improves key phrase extraction. It can be seen from the comparison experiment results that both the global similarity and the local similarity proposed by the present invention play a role. In the calculation of the local similarity, the noise filtering threshold θ proposed by the present invention. In addition, the model proposed by the present invention has made great progress on long texts, which benefits from the fact that the present invention makes full use of the hierarchical structure of the document. The present invention makes the candidate phrases have higher semantic diversity through diversity operations, making the results more acceptable.
[0075] Aiming at the problem that the encoding length of BERT is limited, resulting in that traditional embedding-based methods can only truncate long texts, causing a large amount of semantic loss. According to the document hierarchical structure and human writing habits, the present invention proposes to group the title and abstract as one group and the conclusion as another group and send them into the pre-trained model for embedding twice when facing long texts. This not only saves time and space but also maximally preserves the semantic information of the full text.
[0076] In view of the deficiencies of traditional methods in processing long texts, the present invention creatively proposes to use SimCSE to segment and encode the beginning and conclusion of a document, enabling the model proposed by the present invention to fully learn the information of the document and obtain higher-quality key phrases.
[0077] In view of the problem of preference for long and short phrases caused by semantic space alignment, according to the writing habits of humans, the present invention uses the title and the last sentence of the ending to replace the traditional full-text vector, thereby solving the problem of high scores for long and short phrases.
[0078] The present invention performs a post-processing operation on the candidate phrases. A threshold is set to filter out the candidate phrases that appear frequently in the top 20% in each specific field, avoiding the interference of high-frequency and ineffective phrases, and then improving the semantic diversity of the candidate key phrases by deleting subsets.
[0079] In the local text information modeling, the present invention adopts topic centrality, which can identify the topic information in the full text and can better capture the local topic information compared with boundary centrality. The present invention solves the technical problems of semantic loss, preference for long and short phrases, and insufficient extraction of key information from the main body in the prior art, resulting in low accuracy of key phrase extraction. Brief Description of the Drawings
[0080] Figure 1 It is a schematic diagram of the method steps for extracting keywords by combining document hierarchy and global-local information in Embodiment 1 of the present invention;
[0081] Figure 2 It is a schematic diagram of the overall process for extracting keywords by combining document hierarchy and global-local information in Embodiment 1 of the present invention;
[0082] Figure 3 It is a schematic diagram of the structure similarity vector and similarity in Embodiment 1 of the present invention;
[0083] Figure 4 It is a schematic diagram of the visualization principle of key phrases and full-text vector embedding in Embodiment 1 of the present invention;
[0084] Figure 5 It is a schematic diagram of a test example in Embodiment 3 of the present invention. Detailed Embodiments
[0085] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0086] Example 1
[0087] As Figure 1 and Figure 2 shown, the method for extracting key phrases by combining global and local information of the document hierarchy provided by the present invention includes the following steps:
[0088] S1. Tokenize and perform part-of-speech tagging on the input document, and generate candidate key phrases according to rules;
[0089] In this embodiment, the StanfordCoreNLP tool is used to tokenize and perform part-of-speech tagging on the input document, and candidate key phrases are generated according to rules; in this embodiment, the standard natural language processing tool StanfordCoreNLP is used to tokenize and perform part-of-speech tagging on the document D. After tokenization, the document can be represented as D = {T 1 , T 2 , …, T n}.
[0090] In this embodiment, the standard Stanford CoreNLP tool is used to tokenize and perform part-of-speech tagging on the document D. First, a general stop word list is set to filter out words and symbols that have no meaning. The specific method is that if a word appears in the stop word list, then its part-of-speech tag is marked as 'IN'. Then the rule <NN.*|JJ>*<NN.*> is used, where NN.* represents a noun and JJ represents an adjective. This rule extracts phrases of any length that start with a noun or an adjective and end with a noun. Using this rule, a set of candidate key phrases KP = {KP 0 , KP 1 , …, KP n} can be obtained.
[0091] S2. Send all short documents into BERT for embedding, and select the first 512 and the best 512 words of long documents for embedding and obtain vector representations of the title, the last sentence, and candidate key phrases;
[0092] In this embodiment, the length of the input document is judged. If the length of the document does not exceed the threshold, the BERT model is directly used for embedding to obtain a vector representation. If the length of the document exceeds the threshold, the content within the specified range of the document is input into the SimCSE model for embedding and a vector representation is obtained;
[0093] In this embodiment, after preprocessing, the document D is preprocessed into a tokens set T = {T 1 , T 2 , …, T n} and a candidate key phrase sequence KP = {KP 0 , KP1 , …, KP m}, different from using static vectors for encoding before, in the present invention, BERT is used for vector embedding. This is a very strong pre-trained model that can obtain dynamic vector representations of the context. In this way, vector representations of candidate key phrases and documents can be obtained. Considering the hierarchical structure information of the document, the vector representation of the document here uses the title H title and the last sentence at the end H end .
[0094] In this embodiment, as follows Figure 3 shown, where the parameter symbols are defined as in Table 1 below:
[0095] Table 1: Definition of Parameter Symbols
[0096]
[0097] As Figure 3 shown, input D = {T 1 , T 2 , …, T n} obtained in the first step into the BERT model, insert the CLS token at the beginning and the SEP token at the end, and then perform embedding learning to obtain the vector of each token, that is: {H 1 , H 2 , …, H n} = BERT({T 1 , T 2 , …, T n}). Then, according to the extraction rules in the second step, obtain the vector representation of the candidate key phrases. For candidate key phrases composed of multiple words such as "key phrase extraction", the MaxPooling method is selected to obtain its vector. That is, the candidate key phrase vector set is In this step, the title and the end can be sent into the BERT model to obtain the corresponding vectors H title , H end .
[0098] Since BERT can only encode 512 tokens, existing methods cannot handle long texts well. For most long documents or news articles, the author tends to write key information at the beginning and end of the document. Therefore, for practical operation convenience, the present invention inputs the conclusion and abstract of the article into the BERT model for embedding respectively, so as to obtain more and more diverse candidate key phrases, thus fully mining the document information. However, the research of [LingxiaoWang, 2020] points out that the vector expressions encoded by BERT have anisotropy, which is manifested as uneven distribution. The distribution of low-frequency words is sparse while that of high-frequency words is dense. Therefore, it is impossible to measure the similarity between two sentences. Inspired by the research of [Tianyu Gao, 2021], the present invention uses SimCSE to replace the expression of BERT on long texts. SimCSE is a model proposed by the group of Danqi Chen at Princeton University, aiming to solve the anisotropy problem of the vectors encoded by BERT.
[0099] S3. Use the title, the last sentence and the candidate key phrase vectors to calculate the global similarity;
[0100] In this embodiment, for the global similarity measure, the present invention innovatively uses the document title and the end to evaluate the global similarity of the candidate key phrases, so as to solve the preference for longer candidate key phrases caused by vector space alignment;
[0101] In this embodiment, for each candidate key phrase, use the similarity method to calculate its similarity with the title H title and the last sentence H end . In this embodiment, after the previous steps, the vector representation of each candidate key phrase has been obtained The vector representation of the title H title and the vector representation of the end H end . The more similar the phrase is to the article, the more likely it is to be a key phrase. However, the input sequence lengths of the document and the phrase in the BERT model are not equal, resulting in difficulty in aligning the semantic spaces. Longer phrases have more advantages than single words. Inspired by the document structure, considering that people usually put their core views at the beginning and end of the article, the present invention uses the title and the end to replace the full-text vector, because it can minimize the length gap with the candidate key phrases as much as possible, and using the title and the end to replace the full-text vector will reduce some noise.
[0102] The global similarity of each candidate key phrase i is calculated by the following formula:
[0103]
[0104] where ‖.‖ represents the Manhattan distance, It represents the global similarity between the candidate key phrase i and the entire document. H title It represents the vector of the title, H end It represents the vector of the ending.
[0105] S4. Construct a complete undirected graph, where the nodes are candidate key phrases and the edges are the similarities between the nodes. Then, set an adaptive threshold according to the maximum and minimum values of each document, and update the weight of the edge to weight - threshold. Set the weight of the new edge less than 0 directly to 0. In this way, achieve topic division and calculate local similarity;
[0106] Such as Figure 4 As shown, in this embodiment, for local similarity evaluation, a new topic centrality is adopted to perform topic division and clustering on the candidate key phrases of the full text, fully capturing local significant information. In the present invention, after embedding the key phrases and the full - text vector and visualizing them, the pentagram is the vector embedding of the article, and the nodes with the same filling belong to the same topic.
[0107] In this embodiment, construct a complete undirected graph, where the vertices are candidate key phrases. The initial weight of the edge is the dot - product result between two term vectors. Considering that an article is composed of multiple small topics, a dynamic - threshold method is adopted to filter the noise of irrelevant topics to the candidate key phrases.
[0108] In this embodiment, first construct a complete undirected graph G=(V, E), where the points are That is, the points are candidate key phrases. The edges are E = {e ij}, which represents the weights between candidate key phrases. The traditional graph centrality calculation method is:
[0109]
[0110] Among them, According to what we described above, a document will have multiple local topics, and candidate key phrases will form multiple local topics. Therefore, how to accurately find these small topics is a difficult problem. It can be seen from Figure 4 that the candidate key phrases included in each small topic are clustered together. And when the topic difference is not particularly large, a candidate key phrase may be included in multiple topics. For a candidate key phrase, if it is included in more small topics, it means it is more important. However, phrases in completely opposite topics will cause noise interference to the candidate key phrase. Based on this assumption, the present invention designs a threshold θ to filter the noise.
[0111] θ = min(e ij ) + β×(max(e ij)-min(e ij ))
[0112] Below this threshold θ, we set the weight e of the edge ij = 0, so as to filter out the interference of completely irrelevant phrases. In this way, we rewrite the traditional degree centrality calculation formula as
[0113]
[0114] Here represents the local significance of the candidate key phrase i.
[0115] For most documents, the author tends to write the key information at the beginning of the document. [Florescu, 2017] pointed out that the position bias weight can greatly improve the performance of key phrase extraction. They use the sum of the reciprocal of the position and the words in the document as the weight. For example, if a candidate key phrase appears in the second, fifth, and tenth positions, then its position score To prevent double counting, the present invention only calculates the position where the candidate key phrase appears for the first time as its position score, that is where p 1 is the position where the candidate key phrase i appears for the first time. To prevent the position information from dominating the final score, the softmax function is used to smooth the position score. Therefore, the present invention modifies the position information formula to:
[0116]
[0117] Therefore, after comprehensively considering the position information, the present invention rewrites the local similarity formula of the candidate key phrase i as follows.
[0118]
[0119] The present invention finally uses to measure the local similarity of the candidate key phrase i.
[0120] S5. Sorting algorithm: Combine the position information, global similarity, and local similarity to score and sort the subsequent key phrases;
[0121] In this embodiment, the candidate key phrases are comprehensively evaluated and scored by combining the position information, global similarity, and local similarity, and then ranked according to the scores;
[0122] In this embodiment, a large number of existing documents have proven that for most research papers and news articles, authors tend to write key information at the beginning and end of the document. Therefore, the position information of candidate key phrases is very important. In addition, the more times a phrase appears in an article, the more likely it is to be a key phrase. Considering that the word frequency information has already been used in the local similarity calculation process, to prevent duplicate calculation, therefore, the present invention only records the first position where a phrase appears, and takes the reciprocal of the position of the candidate key phrase as the position score. The global similarity score, local similarity score, and position score of a phrase are comprehensively calculated. Finally, a score list is output. In this embodiment, the present invention combines the global similarity and local similarity of candidate key phrases in the most concise multiplication method, so the final score of a candidate key phrase is:
[0123]
[0124] S6. Post-processing: For the candidate key terms after the sorting algorithm, improve the diversity by deleting subsets in the candidate key phrases and then delete high-frequency common words to overall improve the user experience;
[0125] In this embodiment, a post-processing operation is performed to remove high-frequency common words from the data set to avoid interference from high-frequency invalid words, and then the semantic diversity of candidate key phrases is improved by deleting subsets.
[0126] In this embodiment, a semantic diversity operation is performed on the key phrases, then the high-frequency common phrases are filtered, and finally the top N are selected as the key phrases.
[0127] In this embodiment, considering the diversity of key phrases, more detailed key phrases are used to replace coarser-grained key phrases in an article. Therefore, the present invention selects to use fine-grained key phrases to delete coarser-grained key phrases. For example, if "government policies" and "government", "policies" appear in the candidate key phrase list, then we will delete the two coarser-grained candidate key phrases "government" and "policies", so as to obtain key phrases with more diversity and more in line with human needs.
[0128] A candidate key phrase can be formally represented as KP i ={w 1 ,w 2 ,…,w n}. This method deletes the single-word w 1 ,w 2 ,…,w n in the candidate key phrases, so as to obtain key phrases with more semantic diversity.
[0129] Example 2
[0130] The method for extracting keywords by combining global and local information in the document hierarchy provided by the present invention has been verified in three public datasets, Inspec, DUC2001, and SemEval2010. The results show that the method and device for extracting key phrases by combining global and local information in the document hierarchy proposed by the present invention can effectively achieve document key phrase extraction.
[0131] Experimental results:
[0132] Dataset
[0133] The present invention conducts experiments on three public datasets, namely Inspec [Hulth, 2003], DUC2001 [Wan, 2008], and SemEval2010 [Kim, 2010]. The Inspec dataset contains 2000 abstract documents from scientific journal abstracts. In our experiment, 500 test documents and the version of key phrases annotated by readers are used as GroundTruth. DUC2001 is a collection of 308 long news articles. SemEval2010 contains ACM full-length papers. In this experiment, 100 test documents and a combined set of key phrases annotated by authors and readers are used.
[0134] Experimental results:
[0135] As shown in the following table, like previous work, the present invention selects three metrics, F1@5, F1@10, and F1@15, to evaluate the accuracy of the method (DHSRank+) proposed by the present invention.
[0136]
[0137] It can be seen from the experimental results that the method proposed by the present invention has greatly improved key phrase extraction. By comparing the experimental results, it can be seen that both the global similarity and the local similarity proposed by the present invention play a role. In the calculation of local similarity, the noise filtering threshold θ proposed by the present invention is particularly crucial. In addition, the model proposed by the present invention has made great progress on long texts, which benefits from the fact that the present invention makes full use of the document hierarchy. Especially for the deficiencies of traditional methods in processing long texts, the present invention innovatively proposes to use SimCSE to segment and encode the beginning and conclusion of the document, enabling the model proposed by the present invention to fully learn the information of the document and obtain higher-quality key phrases. Finally, the present invention makes the candidate key phrases have higher semantic diversity through diversity operations, making the results more acceptable.
[0138] Example 3
[0139] Example demonstration:
[0140] As Figure 5 shown, an example in DUC2001
[0141] DUC2001 is a dataset from news articles. The correct key phrases are underlined with a solid line. The black bold text represents the standard key phrases, and the text marked with a dashed underline represents the phrases extracted by our model.
[0142] We can see that the gold standard corresponds to the various themes of the article. Our model has extracted many correct phrases that are the same as the standard key phrases, and has extracted the phrase "net income" which has the same semantics as "lower net income" in the standard key phrases.
[0143] It is worth mentioning that our model focuses on the boundaries of the document, and most of the extracted phrases are located at the beginning and end of the document, which proves the effectiveness of the title + end as the global vector proposed by us. From the figure, we can also find that the incorrect phrases are highly correlated with the various small themes of the document, which proves the effectiveness of our topic-aware centrality. This example shows that the joint modeling of global and local contexts can improve the performance of key phrase extraction, and our model truly captures local and global information.
[0144] In summary, the method proposed by the present invention has greatly improved key phrase extraction. It can be seen from the comparison experiment results that both the global similarity and the local similarity proposed by the present invention have played a role. In the calculation of local similarity, the noise filtering threshold θ proposed by the present invention. In addition, the model proposed by the present invention has made great progress on long texts, which benefits from the fact that the present invention makes full use of the hierarchical structure of the document. The present invention makes the candidate key phrases have higher semantic diversity through diversity operations, making the results more acceptable.
[0145] Aiming at the problem that the encoding length of BERT is limited, resulting in that traditional embedding-based methods can only truncate long texts, leading to a large amount of semantic loss, the present invention, based on the document hierarchical structure and human writing habits, proposes to group the title and abstract as one group and the conclusion as another group and send them into the pre-trained model for embedding twice when facing long texts, which not only saves time and space, but also maximally preserves the semantic information of the full text.
[0146] Aiming at the deficiencies of traditional methods in processing long texts, the present invention creatively proposes to use SimCSE to segment and encode the beginning and conclusion of the document, so that the model proposed by the present invention can fully learn the information of the document and obtain higher-quality key phrases.
[0147] The present invention addresses the problem of preference for long and short phrases caused by semantic space alignment. According to human writing habits, it uses the title and the last sentence at the end to replace the traditional full-text vector, thereby solving the problem of high scores for long and short phrases.
[0148] The present invention performs a post-processing operation on the candidate key phrases. A threshold is set to filter out the candidate key phrases that appear frequently in the top 20% in each specific field, avoiding the interference of high-frequency invalid phrases, and then the semantic diversity of the candidate key phrases is improved by deleting subsets.
[0149] In the local text information modeling of the present invention, topic centrality is adopted, which can identify the topic information in the full text and can capture the local topic information better than boundary centrality. The present invention solves the technical problems of semantic loss, preference for long and short phrases, and insufficient extraction of subject information in the prior art, resulting in low accuracy of keyword extraction.
[0150] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for extracting key phrases by combining global and local information in a document hierarchy, characterized in that, the method includes: S1. Use the StanfordCoreNLP tool to tokenize and perform part-of-speech tagging on the input document, and perform NP chunking according to preset extraction rules to generate a set of candidate key phrases; S2. Determine whether the length of the input document is less than or equal to a preset document length threshold. If so, use the BERT model to embed and process the input document to obtain a vector representation. If not, obtain the specified range content of the input document according to a preset range, and input the specified range content into the SimCSE model to perform embedding to obtain the vector representation, title vector, and ending vector of the candidate key phrases; S3. Process the title vector and the ending vector to perform global similarity measurement on the candidate key phrases, and obtain a global similarity accordingly; S4. Use topic centrality to perform topic division and clustering on the candidate key phrases in the full text of the input document according to a preset logic, and obtain a local similarity through local similarity evaluation, where the step S4 further includes: S41. Use the candidate key phrases as nodes and the similarity between the nodes as edges to construct a complete undirected graph; S42. Set an adaptive noise filtering threshold according to the maximum and minimum values of each input document, where the step S42 includes: S421. Use a graph centrality calculation method to process the candidate key phrase i according to the following logic: Among them, S422. Set the adaptive noise filtering threshold θ according to the following logic; θ = min(e ij ) + β × (max(e ij ) - min(e ij )) S43. Update the weights of the edges according to the adaptive noise filtering threshold to obtain local significance data, and obtain updated edges according to the local significance data, where the step S43 includes: S431. Use the following logic to process and obtain the local significance data: Among them, represents the local significance of the candidate phrase i; S432. Obtain the updated edges according to the local significance data. When the weight of the updated edge is less than 0, set the weight of the updated edge to 0; S44. Obtain the position information of the candidate key phrases in the full text of the input document, where the step S44 includes: S441. Calculate the first occurrence position of the candidate key phrase in the input document according to the following logic as the candidate key phrase position score: where p 1 is the position where the candidate term i first appears; S442. Smoothly process the candidate key phrase position score using the softmax function to process and obtain the position information according to the following logic: S45. Calculate the local similarity based on the position information. In step S45, the position information is processed using the following logic to obtain the local similarity of the candidate key phrase i S5. Combine and process the position information, the global similarity, and the local similarity to comprehensively evaluate and score the candidate key phrases, and sort the candidate key phrases accordingly to obtain key phrase ranking data; S6. Obtain a sorted dataset of candidate key phrases according to the key phrase ranking data, perform post-processing operations on the candidate key phrases, delete subsets of the candidate key phrase set to obtain semantically diverse key phrases, obtain lexical frequency data, and remove high-frequency general phrases from the sorted dataset of candidate key phrases to filter out interference from high-frequency invalid phrases.
2. The method for extracting key phrases by combining global and local information of a document hierarchy according to claim 1, characterized in that, the step S2 includes: S21. Insert a CLS token at the start position of the input document and a SEP token at the end position using a BERT model; S22. Embed and learn the input document to obtain a vector for each token: {H 1 ,H 2 ,…,H n} = BERT({T 1 ,T 2 ,…,T n}); S23. Then obtain the vector representation of the candidate key phrases according to the preset extraction rules to obtain the candidate phrase vector set: S24. Feed the title and the ending of the input document into the BERT model to obtain a title vector H title and an ending vector H end; S25. Input the conclusion and abstract of the input document into the BERT model for embedding operations to obtain the vector representation; S26. Use the SimCSE model to perform an expression on the long text of the input document.
3. The method for extracting key phrases by combining global and local information of a document hierarchy according to claim 1, characterized in that, In the step S3, the title vector H is processed according to the following logic title and the ending vector H end , so as to obtain the global similarity of each candidate key phrase i where, ‖.‖ represents the Manhattan distance, represents the global similarity between the candidate phrase i and the entire document.
4. The method for extracting key phrases by combining global and local information of a document hierarchy according to claim 1, characterized in that, the step S5 includes: S51. Use the following logic to perform multiplicative comprehensive processing on the global similarity and the local similarity of the candidate key phrases to obtain a candidate key phrase score: S52. Sort the candidate key phrases according to the candidate key phrase score to obtain the key phrase ranking data.
5. The method for extracting key phrases by combining global and local information of a document hierarchy according to claim 1, characterized in that, in the step S6, the coarse-grained key phrases are deleted according to the fine-grained key phrases to obtain the semantically diverse key phrases.
6. A system for extracting key phrases by combining global and local information of a document hierarchy, which is used to execute the method for extracting key phrases by combining global and local information of a document hierarchy according to any one of the preceding claims 1 to 5, characterized in that, the system includes: A candidate phrase generation module that uses the StandfordCoreNLP tool to tokenize and perform part-of-speech tagging on the input document, and performs NP chunking according to preset extraction rules to generate a candidate key phrase set; A BERT model embedding module that is used to determine whether the length of the input document is less than or equal to a preset document length threshold. If so, use the BERT model to embed and process the input document to obtain a vector representation. If not, obtain the specified range content of the input document according to a preset range, and input the specified range content into the BERT model to perform embedding to obtain the vector representation of the candidate key phrases, the title vector, and the ending vector. The BERT model embedding module is connected to the candidate phrase generation module; A global similarity measurement module that is used to process the title vector and the ending vector to perform global similarity measurement on the candidate key phrases to obtain a global similarity. The global similarity measurement module is connected to the BERT model embedding module; A local similarity evaluation module, which uses topic centrality to perform topic division and clustering on the candidate key phrases in the full text of the input document according to preset logic, and obtains local similarity based on local similarity evaluation. The local similarity evaluation module is connected to the candidate phrase generation module. Among them, the local similarity evaluation module further includes: An undirected graph construction module, which uses the candidate key phrases as nodes and the similarity between the nodes as edges to construct a complete undirected graph; A noise filtering threshold setting module, which is used to set an adaptive noise filtering threshold according to the maximum and minimum values of each input document; A noise filtering module, which is used to update the weights of the edges according to the adaptive noise filtering threshold to obtain local significance data, and obtain updated edges according to the local significance data. The noise filtering module is connected to the undirected graph construction module and the noise filtering threshold setting module; A position acquisition module, which is used to obtain the position information of the candidate key phrases in the full text of the input document according to the new complete undirected graph; A local similarity calculation module, which is used to calculate the local similarity according to the position information. The local similarity calculation module is connected to the position acquisition module; A key phrase ranking module, which is used to comprehensively evaluate and score the candidate key phrases by combining and processing the position information, the global similarity, and the local similarity, and sort the candidate key phrases accordingly to obtain key phrase ranking data. The key phrase ranking module is connected to the global similarity measurement module and the local similarity evaluation module; A post-processing module, which is used to obtain a candidate key phrase sorting data set according to the key phrase ranking data, perform post-processing operations on the candidate key phrases, delete subsets of the candidate key phrase set to obtain semantically diverse key phrases, obtain phrase frequency data, and remove high-frequency common phrases from the candidate key phrase sorting data set to filter out interference from high-frequency invalid phrases. The post-processing module is connected to the key phrase ranking module.
Citation Information
Patent Citations
Keyword extraction method, verbal skill scoring method and verbal skill recommendation method
CN111160017A
Theme extraction method and device oriented to science and technology requirements and storage medium
CN113255340A
Automatic key phrase extracting method and system for English literatures
CN106066866A
Method and system for extracting Chinese key phrases in scientific and technological innovation field by utilizing semantic features
CN113221559A