A keyword extraction method based on graph model and word embedding model
By applying the keyword extraction method based on graph model and word embedding model in news text, the problem of insufficient keyword extraction in the prior art and the neglected verb function is solved, and more efficient and comprehensive keyword extraction is achieved, and the accuracy of public opinion analysis is improved.
Patent Information
- Application Number
- CN202210606979.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-05-31
AI Technical Summary
The existing keyword extraction technology cannot fully display information in news texts, ignore important parts, and the role of verbs is ignored, resulting in one-sided extraction results.
The keyword extraction method based on the graph model and word embedding model is adopted, and the final keyword distribution is obtained by cleaning news texts and combining pre-trained word embedding models and graph models.
It improves the accuracy of keyword extraction of news texts, covers the main information of news texts more comprehensively, saves manual review time, and can analyze public opinion events more accurately.
Smart Images

Figure CN115034216B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and specifically, relates to a keyword extraction method based on a graph model and a word embedding model. Background Art
[0002] Keyword extraction is a method of extracting key information from text. It is often used in retrieval systems, manual text screening, manual text category annotation and other scenarios. In recent years, keyword extraction has also been applied to public opinion monitoring systems.
[0003] Ten years ago, paper media may have been very competitive in the media industry, but nowadays, mobile phones are in the hands of every citizen. Compared with paper media, electronic media that allows users to quickly flip through pages is now in full swing. At this time, countless electronic articles are overwhelming, and it is obviously unrealistic to read these endless articles word by word manually. Even if we let search engines filter articles for us, we still need to check the content of the articles, and one of the technologies required is keyword extraction.
[0004] On the other hand, with the development of many online public opinion events, a customized public opinion monitoring and analysis system has become an urgent need for major companies and enterprises. Extracting important information from online texts and presenting it to workers or screening online texts coincides with the information that keywords can provide.
[0005] However, the current keyword extraction technology is not very mature and there are still many areas that need improvement:
[0006] First, the existing keyword extraction technology obtains relatively short words, which cannot fully display all the information and may even ignore important parts;
[0007] Second, when extracting keywords, the important role of verbs is ignored in some scenarios, especially in the field of public opinion news, where verbs often play a critical role and set the tone for the nature of an event;
[0008] 3. There are many methods for extracting keywords, but the current methods mainly use a single method for extraction, and fewer methods will combine multiple methods for extraction, resulting in the extracted keywords being one-sided. Summary of the invention
[0009] In order to improve the accuracy of keyword extraction in news texts, further improve the accuracy of content retrieval of the public opinion analysis system when analyzing news texts, more comprehensively cover the main information of news texts, and save time for manual review, the present invention proposes a keyword extraction method based on a graph model and a word embedding model.
[0010] The present invention is achieved through the following technical solutions:
[0011] A keyword extraction method based on graph model and word embedding model:
[0012] The method specifically comprises the following steps:
[0013] Step 1: Clean the news text and remove invalid information;
[0014] Step 2: Process the news text cleaned in step 1 to obtain the candidate keywords (groups), positions and word frequency information;
[0015] Step 3: Use the pre-trained word embedding model to embed and calculate the text obtained in steps 1 and 2, obtain the vector representation of each word segment and the entire article, perform similarity calculation, and obtain keyword distribution 1;
[0016] Step 4: Use the text information from steps 1 and 2 and the text vector representation from step 3 and apply them to the graph model to obtain keyword distribution 2;
[0017] Step 5: Combine keyword distribution 1 and keyword distribution 2 to obtain the final keyword distribution and complete the acquisition of news text keywords.
[0018] Furthermore, in step 1,
[0019] The invalid information includes traditional Chinese characters in news text, formatting symbols of website layout, and fixed text at the beginning or end of an article.
[0020] Furthermore, in step 1,
[0021] Step 1.1: Use regular expressions to clean the obtained news text, retain Chinese, English and numeric characters, and remove useless web links, emoticons, spaces and non-Chinese characters;
[0022] Step 1.2: Use the library function in the Python programming language to simplify the source text and convert the traditional Chinese characters into simplified Chinese. If the original text does not contain traditional Chinese characters, skip this step.
[0023] Step 1.3: Use regular expressions to remove fixed representations.
[0024] Furthermore, in step 2,
[0025] Step 2.1: Use the word segmentation tool to segment the text obtained in step 1 to obtain the word segmentation result;
[0026] Step 2.2: Use a part-of-speech tagging tool to tag the above word segmentation results to obtain the part of speech of each word segmentation;
[0027] Step 2.3: Rules for constructing a grammar parse tree: retain participles such as people, places, verbs, and nouns, and if there are consecutive nouns or adjectives and nouns, combine them together to form a candidate keyword group;
[0028] Step 2.4: Use a grammar parsing tool to process the above rules and part-of-speech tagging results to obtain candidate keywords (groups), and at the same time obtain the position offset of each candidate keyword (group) relative to the first character of the source text, that is, the position information;
[0029] Step 2.5: Count the frequency information of the keywords (groups) selected in step 2.4 appearing in the source text.
[0030] Furthermore, in step 3,
[0031] Step 3.1: Input the word segmentation results obtained in step 2.1 into the pre-trained model ELMO to obtain the word embedding representation of each word in each layer
[0032] Where l∈{0,1,2} represents the first, second, and third LSTM layers of EMLO respectively; i∈[0,N] represents the i-th position in the article segmentation result. It represents the word embedding representation of the i-th word segmentation result of the article in the l-th layer of the ELMO model, and N represents the number of word segmentation results of the article;
[0033] Step 3.2: The three layers of ELMO have different weights. The embedding of each word is obtained according to the weights and the word embedding representation obtained in step 3.1. i , the formula is as follows:
[0034]
[0035] Step 3.3: According to step 2.3 and step 2.4, obtain the candidate keyword (group) and position information, add the word embedding representation vectors of the word segments involved in the candidate keyword to obtain the candidate keyword (group) representation KeyPhrase i , when adding vectors, the relative position information of each word in the current candidate keyword group is considered. The specific fusion formula is as follows:
[0036]
[0037] Where m means that the keyword group to be selected consists of m segmentation results. i,j Represents the embedding representation of the jth word segmentation result in the i-th candidate keyword group;
[0038] Step 3.4: Based on the frequency information of the candidate keywords (groups) obtained in step 2.5 and the representation of each word obtained in step 3.3, calculate the embedded vector representation of the article. The calculation formula is as follows:
[0039]
[0040] Among them, Fre i It represents the frequency of occurrence of the i-th candidate keyword (group), and N represents the number of article segmentation results.
[0041] Further,
[0042] Step 3.5: Calculate the cosine similarity based on the candidate keyword (group) representation and article representation obtained in step 3.3 and step 3.4. The formula is as follows:
[0043] similarity i =cos(docEmbedding,KeyPhrase i )
[0044] Step 3.6: Use the frequency dictionary that comes with the Jieba word segmentation function to correct the result of step 3.5. The formula is as follows:
[0045]
[0046] Among them, JiebaFre i represents the default frequency of the i-th candidate keyword (group) in the Jieba word segmentation vocabulary;
[0047] Step 3.7: Correct the result of step 3.6 in combination with the position of each candidate keyword (group), the formula is as follows:
[0048]
[0049] where pos i Indicates the first occurrence position of each candidate keyword (group) in the original text;
[0050] Step 3.8: Similarity to step 3.7 i Combining them, we get keyword distribution 1:
[0051] distribution1={similarity”0,…,similarity” N}.
[0052] Furthermore, in step 4,
[0053] Step 4.1: Follow the process of obtaining dicEmbedding in step 3 to obtain the title representation of each news, i.e. titleEmbedding;
[0054] Step 4.2: Construct a graph model TextRank model, in which nodes represent candidate keywords (groups), node weights nodeWeight represent the importance of candidate keywords (groups), and edges between nodes are established only when two nodes appear simultaneously in a fixed-size window in the original text;
[0055] Step 4.3: Initialize the node weights of the graph model according to the vector representation of the candidate keyword (group) obtained in step 3.3, and calculate the initial node weights according to the following formula:
[0056]
[0057] Where N represents the number of all keywords (groups) to be selected, and α∈[0,1] is the adjustment factor;
[0058] Step 4.4: Initialize the weights of the edges in the graph model based on the vector representation of the candidate keywords (groups) obtained in step 3.3, and calculate the initial weights of the edges according to the following formula:
[0059]
[0060] Among them, Fre i,j represents the frequency of nodes i and j appearing in a fixed-size window, β∈[0,1] is the adjustment factor;
[0061] Step 4.5: Use the node weights and edge weights redefined above to calculate TextRank. After the graph model converges, the converged weight values of the nodes are obtained, and the keyword distribution 2 is obtained:
[0062] distribution2={nodeWeight0,…,nodeWeight N}.
[0063] Furthermore, in step 5,
[0064] Step 5.1: Obtain the keyword distribution 1 and keyword distribution 2 of step 3.8 and step 4.5, and obtain the similarity value and importance corresponding to each keyword in the keyword distribution 1 and keyword distribution 2;
[0065] Step 5.2: According to the result of step 5.1, the distribution of the final candidate keywords (groups) is calculated as follows:
[0066] score i =γ·similarity” i +(1-γ)·nodeWeight i
[0067] FinalDistribution={score0,…,score N}
[0068] Step 5.3: Select the top K candidate keywords (groups) with the highest values from FinalDistribution as the final keyword results.
[0069] An electronic device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.
[0070] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the steps of any of the above methods are implemented.
[0071] Beneficial effects of the present invention
[0072] The present invention obtains information such as keywords to be extracted and word frequency by cleaning the original news text, models keywords using a graph model and a pre-trained model, and finally integrates the results of the two methods to obtain the final keywords required;
[0073] The method of the present invention can be applied in search engines to help the system better handle retrieval needs; or it can be applied in public opinion monitoring systems to help users analyze key information of current public opinion events; or it can be directly applied to news texts on major websites to pre-acquire keywords for users and place them on the page, which helps users screen target articles in a short time and quickly lock in required information, thereby improving user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 is a system flow chart of the present invention;
[0075] Figure 2 A flowchart for cleaning text in the present invention;
[0076] Figure 3 A model diagram for text word embedding in the present invention;
[0077] Figure 4 This is an example diagram of the present invention. DETAILED DESCRIPTION
[0078] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0079] Combination Figures 1 to 4 .
[0080] A keyword extraction method based on graph model and word embedding model:
[0081] The method specifically comprises the following steps:
[0082] Step 1: Clean the news text and remove invalid information;
[0083] Step 2: Process the news text cleaned in step 1 to obtain the candidate keywords (groups), positions and word frequency information;
[0084] Step 3: Use the pre-trained word embedding model to embed and calculate the text obtained in steps 1 and 2, obtain the vector representation of each word segment and the entire article, perform similarity calculation, and obtain keyword distribution 1;
[0085] Step 4: Use the text information from steps 1 and 2 and the text vector representation from step 3 and apply them to the graph model to obtain keyword distribution 2;
[0086] Step 5: Combine keyword distribution 1 and keyword distribution 2 to obtain the final keyword distribution and complete the acquisition of news text keywords.
[0087] In step 1,
[0088] In real life, after news text is obtained, it will contain some invalid information or information that is relatively unimportant to the source text, but this information will have a certain impact on keyword extraction. This includes traditional Chinese characters in news text, formatting symbols for website layout, and some fixed text that appears at the beginning or end of the article, such as an introduction to who wrote the article or an introduction to the operation of a WeChat public account, and some text descriptions at the end of WeChat public account articles such as "Take a look - Search".
[0089] The invalid information includes traditional Chinese characters in news texts, formatting symbols for website layout, and fixed text at the beginning or end of articles.
[0090] This step is one of the pre-steps, which aims to clean the acquired news text. After cleaning, a clean source text is obtained, which can greatly facilitate subsequent related processing.
[0091] In step 1,
[0092] Step 1.1: Use regular expressions to clean the obtained news text, retain Chinese, English and numeric characters, and remove useless web links, emoticons, spaces and non-Chinese characters;
[0093] Step 1.2: Use the library function in the Python programming language to simplify the source text and convert the traditional Chinese characters into simplified Chinese. If the original text does not contain traditional Chinese characters, skip this step.
[0094] Step 1.3: Use regular expressions to remove fixed expressions, such as "Editor in Charge: XXX" or WeChat official account introduction text and "Take a look - Search" text;
[0095] In step 2,
[0096] This step is one of the pre-steps. The news text obtained in step 1 with invalid information removed is processed to obtain the candidate keywords (groups), as well as the corresponding position information and the number of occurrences of each candidate keyword or candidate keyword group in the source text. This is applied in steps 3 and 4.
[0097] Step 2.1: Use the word segmentation tool to segment the text obtained in step 1 to obtain the word segmentation result;
[0098] Step 2.2: Use a part-of-speech tagging tool to tag the above word segmentation results to obtain the part of speech of each word segmentation;
[0099] Step 2.3: Rules for constructing a grammar parsing tree: retain the participles such as people, places, verbs, and nouns (because the keyword extraction is for news text, the people, places, and actions of the news need to be considered here), and if there are consecutive nouns or adjectives and nouns, combine them together to form a candidate keyword group;
[0100] Step 2.4: Use a grammar parsing tool to process the above rules and part-of-speech tagging results to obtain candidate keywords (groups), and at the same time obtain the position offset of each candidate keyword (group) relative to the first character of the source text, that is, the position information;
[0101] Step 2.5: Count the frequency information of the keywords (groups) selected in step 2.4 appearing in the source text.
[0102] In step 3,
[0103] This step is one of the specific keyword extraction steps. The news text obtained in step 1 with invalid information removed is segmented and then input into the pre-trained model to obtain a vector representation of each segmented word.
[0104] Combined with the candidate phrases, positions and word frequency information obtained in step 2, the vector representation of each candidate keyword or candidate keyword phrase is obtained; at the same time, the whole article is segmented and the vector representation of each segmentation result is obtained, and then the vector representation of the whole article is obtained;
[0105] The similarity between the candidate keyword (group) vector and the article vector is calculated, and the final keyword distribution 1 is obtained by combining the position information and frequency information of the candidate keyword (group).
[0106] Step 3.1: Input the word segmentation results obtained in step 2.1 into the pre-trained model ELMO to obtain the word embedding representation of each word in each layer
[0107] Where l∈{0,1,2} represents the first, second, and third LSTM layers of EMLO respectively; i∈[0,N] represents the i-th position of the article segmentation result. It represents the word embedding representation of the i-th word segmentation result of the article in the l-th layer of the ELMO model, and N represents the number of word segmentation results of the article;
[0108] Step 3.2: The three layers of ELMO have different weights. The embedding of each word is obtained according to the weights and the word embedding representation obtained in step 3.1. i , the formula is as follows:
[0109]
[0110] Step 3.3: According to step 2.3 and step 2.4, obtain the candidate keyword (group) and position information, add the word embedding representation vectors of the word segments involved in the candidate keyword to obtain the candidate keyword (group) representation KeyPhrase i It should be noted that the candidate keyword phrase is taken into account here, so when adding the vectors, the relative position information of each word in the current candidate keyword phrase is considered. The specific fusion formula is as follows:
[0111]
[0112] Where m means that the keyword group to be selected consists of m segmentation results. i,j Represents the embedding representation of the jth word segmentation result in the i-th candidate keyword group;
[0113] Step 3.4: Based on the frequency information of the candidate keywords (groups) obtained in step 2.5 and the representation of each word obtained in step 3.3, calculate the embedded vector representation of the article. The calculation formula is as follows:
[0114]
[0115]
[0116] Among them, Fre i It represents the frequency of occurrence of the i-th candidate keyword (group), and N is the number of article segmentation results.
[0117] Step 3.5: Calculate the cosine similarity based on the candidate keyword (group) representation and article representation obtained in step 3.3 and step 3.4. The formula is as follows:
[0118] similarity i =cos(docEmbedding,KeyPhrase i )
[0119] Step 3.6: Use the frequency dictionary that comes with the Jieba word segmentation function to correct the result of step 3.5. The formula is as follows:
[0120]
[0121] Among them, JiebaFre i represents the default frequency of the i-th candidate keyword (group) in the Jieba word segmentation vocabulary;
[0122] Step 3.7: Correct the result of step 3.6 in combination with the position of each candidate keyword (group), the formula is as follows:
[0123]
[0124] where pos i Indicates the first occurrence position of each candidate keyword (group) in the original text;
[0125] Step 3.8: Similarity to step 3.7 i Combining them, we get keyword distribution 1:
[0126] distribution1={similarity”0,…,similarity” N}.
[0127] In step 4,
[0128] This step is one of the specific steps for extracting keywords. Following step 3, we obtain the vector representation of all the word segmentation results of the article and the vector representation of the keywords (groups) to be selected; the other step is to obtain the vector representation of all the word segmentation results of the article title, and then obtain a vector representation of the article title. Next, we use the graph model to iteratively update and calculate the importance of each keyword (group) to be selected, and finally obtain the keyword distribution 2 based on these importances.
[0129] Step 4.1: Follow the process of obtaining dicEmbedding in step 3 to obtain the title representation of each news, i.e. titleEmbedding;
[0130] Step 4.2: Construct a graph model TextRank model, in which nodes represent candidate keywords (groups), node weights nodeWeight represent the importance of candidate keywords (groups), and edges between nodes are established only when two nodes appear simultaneously in a fixed-size window in the original text;
[0131] Step 4.3: Initialize the node weights of the graph model according to the vector representation of the candidate keyword (group) obtained in step 3.3, and calculate the initial node weights according to the following formula:
[0132]
[0133] Where N represents the number of all keywords (groups) to be selected, and α∈[0,1] is the adjustment factor;
[0134] Step 4.4: Initialize the weights of the edges in the graph model based on the vector representation of the candidate keywords (groups) obtained in step 3.3, and calculate the initial weights of the edges according to the following formula:
[0135]
[0136] Among them, Fre i,j represents the frequency of nodes i and j appearing in a fixed-size window, β∈[0,1] is the adjustment factor;
[0137] Step 4.5: Use the node weights and edge weights redefined above to calculate TextRank. After the graph model converges, the converged weight values of the nodes are obtained, and the keyword distribution 2 is obtained:
[0138] distribution2={nodeWeight0,…,nodeWeight N}.
[0139] In step 5,
[0140] This step is to merge the keyword distributions of step 3 and step 4, and select the first K as the extraction results of article keywords according to actual needs.
[0141] Step 5.1: Obtain the keyword distribution 1 and keyword distribution 2 of step 3.8 and step 4.5, and obtain the similarity value and importance corresponding to each keyword in the keyword distribution 1 and keyword distribution 2;
[0142] Step 5.2: According to the result of step 5.1, the distribution of the final candidate keywords (groups) is calculated as follows:
[0143] score i =γ·similarity” i +(1-γ)·nodeWeight i
[0144] FinalDistribution={score0,…,score N}
[0145] Step 5.3: Select the top K candidate keywords (groups) with the highest values from FinalDistribution as the final keyword results.
[0146] An electronic device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.
[0147] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the steps of any of the above methods are implemented.
[0148] Example
[0149] According to the above steps, a simple automatic keyword extraction module can be implemented in the present invention. The module can be embedded in any existing system to achieve a plug-and-play effect. The specific beneficial effects of the invention are as follows:
[0150] This example follows Figure 1 The system built using the present invention is divided into two parts: data acquisition part and algorithm keyword extraction part. The data acquisition part mainly obtains news text from the Internet and stores it in a database or file; the algorithm keyword extraction part mainly includes data cleaning, embedding the selected keywords (groups) into the vector space, and calculating the distribution of the selected keywords (groups), and finally outputs the keywords (groups) of the target news text.
[0151] Figure 4These are news examples that have been crawled and cleaned in the system and keyword (group) results extracted by the system.
[0152] After the system implemented by the present invention is started, the pre-trained model ELMO is first loaded into the memory; then the crawler module is started to collect news texts from the network according to user input;
[0153] The crawler process stores the crawled news text in the system database; at the same time, another process sequentially extracts news text from the system database, uses the algorithm module to extract keywords and stores them in the database;
[0154] When an exception occurs in the above process, the processes where the crawler module and the algorithm module are located are terminated and the system is exited.
[0155] The final actual operation results of the invention can be seen Figure 4 As shown. According to the keyword extraction effect in the figure, it can be seen that the news keyword extraction implemented by the present invention does not only extract a single word but also a phrase, covering information more comprehensively - such as "loan fraud" in the example; at the same time, for the news field, the main characters involved, the place where the event occurred, and related verbs are extracted to describe what happened. Taking into account the limitations of a single algorithm, the present invention combines the graph model, the word embedding model and the statistical model to more comprehensively consider the distribution of keywords.
[0156] The above is a detailed introduction to the keyword extraction method based on the graph model and the word embedding model proposed in the present invention, and the principles and implementation methods of the present invention are explained. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A keyword extraction method based on a graph model and a word embedding model, characterized by: The method specifically comprises the following steps: Step 1: Clean the news text and remove invalid information; Step 2: Process the news text cleaned in step 1 to obtain the candidate keywords (groups), positions and word frequency information; Step 2.1: Use the word segmentation tool to segment the text obtained in step 1 to obtain the word segmentation result; Step 2.2: Use a part-of-speech tagging tool to tag the above word segmentation results to obtain the part of speech of each word segmentation; Step 2.3: Rules for constructing a grammar parse tree: retain participles such as people, places, verbs, and nouns, and if there are consecutive nouns or adjectives and nouns, combine them together to form a candidate keyword group; Step 2.4: Use a grammar parsing tool to process the above rules and part-of-speech tagging results to obtain candidate keywords (groups), and at the same time obtain the position offset of each candidate keyword (group) relative to the first character of the source text, that is, the position information; Step 2.5: Count the frequency information of the keywords (groups) selected in step 2.4 appearing in the source text; Step 3: Use the pre-trained word embedding model to embed and calculate the text obtained in steps 1 and 2, obtain the vector representation of each word segment and the entire article, perform similarity calculation, and obtain keyword distribution 1; Step 3.1: Input the word segmentation results obtained in step 2.1 into the pre-trained model ELMO to obtain the word embedding representation of each word in each layer Where l∈{0,1,2} represents the first, second, and third LSTM layers of EMLO respectively; i∈[0,N] represents the i-th position of the article segmentation result. It represents the word embedding representation of the i-th word segmentation result of the article in the l-th layer of the ELMO model, and N represents the number of word segmentation results of the article; Step 3.2: The three layers of ELMO have different weights. The embedding of each word is obtained according to the weights and the word embedding representation obtained in step 3.
1. i , the formula is as follows: Step 3.3: According to step 2.3 and step 2.4, obtain the candidate keyword (group) and position information, add the word embedding representation vectors of the word segments involved in the candidate keyword to obtain the candidate keyword (group) representation KeyPhrase i , when adding vectors, the relative position information of each word in the current candidate keyword group is considered. The specific fusion formula is as follows: Where m means that the keyword group to be selected consists of m word segmentation results. i,j Represents the embedding representation of the jth word segmentation result in the i-th candidate keyword group; Step 3.4: Based on the frequency information of the candidate keywords (groups) obtained in step 2.5 and the representation of each word obtained in step 3.3, calculate the embedded vector representation of the article. The calculation formula is as follows: Among them, Fre i represents the frequency of occurrence of the i-th candidate keyword (group), and N represents the number of article segmentation results; Step 3.5: Calculate the cosine similarity based on the candidate keyword (group) representation and article representation obtained in step 3.3 and step 3.
4. The formula is as follows: similarity i =cos(docEmbedding,KeyPhrase i ) Step 3.6: Use the frequency dictionary that comes with the Jieba word segmentation function to correct the result of step 3.
5. The formula is as follows: Among them, JiebaFre i represents the default frequency of the i-th candidate keyword (group) in the Jieba word list; Step 3.7: Correct the result of step 3.6 in combination with the position of each candidate keyword (group), the formula is as follows: where pos i Indicates the first occurrence position of each candidate keyword (group) in the original text; Step 3.8: Add the similarity of step 3.7 to i Combining them, we get keyword distribution 1: distribution1={similarity″0,…,similarity″ N } Step 4: Use the text information from steps 1 and 2 and the text vector representation from step 3 and apply them to the graph model to obtain keyword distribution 2; Step 4.1: Follow the process of obtaining dicEmbedding in step 3 to obtain the title representation of each news, i.e. titleEmbedding; Step 4.2: Construct a graph model TextRank model, in which nodes represent candidate keywords (groups), node weights nodeWeight represent the importance of candidate keywords (groups), and edges between nodes are established only when two nodes appear simultaneously in a fixed-size window in the original text; Step 4.3: Initialize the node weights of the graph model according to the vector representation of the candidate keyword (group) obtained in step 3.3, and calculate the initial node weights according to the following formula: Where N represents the number of all keywords (groups) to be selected, and α∈[0,1] is the adjustment factor; Step 4.4: Initialize the weights of the edges in the graph model based on the vector representation of the candidate keywords (groups) obtained in step 3.3, and calculate the initial weights of the edges according to the following formula: Among them, Fre i,j represents the frequency of nodes i and j appearing in a fixed-size window, β∈[0,1] is the adjustment factor; Step 4.5: Use the node weights and edge weights redefined above to calculate TextRank. After the graph model converges, the converged weight values of the nodes are obtained, and the keyword distribution 2 is obtained: distribution2={nodeWeight0,…,nodeWeight N } Step 5: Combine keyword distribution 1 and keyword distribution 2 to obtain the final keyword distribution and complete the acquisition of news text keywords; Step 5.1: Obtain the keyword distribution 1 and keyword distribution 2 of step 3.8 and step 4.5, and obtain the similarity value and importance corresponding to each keyword in the keyword distribution 1 and keyword distribution 2; Step 5.2: According to the result of step 5.1, the distribution of the final candidate keywords (groups) is calculated as follows: scorei=γ·similarity″ i +(1-γ)·nodeWeight i FinalDistribution={score0,…,score N } Step 5.3: Select the top K candidate keywords (groups) with the highest values from FinalDistribution as the final keyword results.
2. The method according to claim 1, characterized in that: In step 1, The invalid information includes traditional Chinese characters in news text, formatting symbols of website layout, and fixed text at the beginning or end of an article.
3. The method according to claim 2, characterized in that: In step 1, Step 1.1: Use regular expressions to clean the obtained news text, retain Chinese, English and numeric characters, and remove useless web links, emoticons, spaces and non-Chinese characters; Step 1.2: Use the library function in the Python programming language to simplify the source text and convert the traditional Chinese characters into simplified Chinese. If the original text does not contain traditional Chinese characters, skip this step. Step 1.3: Use regular expressions to remove fixed representations.
4. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 3 are implemented.
5. A computer-readable storage medium for storing computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Patent literature similarity measurement method based on ontology
CN107247780A
Unsupervised keyword extraction method based on Embedded technology
CN110851570A