Target text keyword extraction method and system under specific category
By introducing antonym clustering and category field filtering in text keyword extraction, a keyword classification list for each category field is constructed, which solves the problem of insufficient text keyword extraction accuracy and achieves more efficient and accurate text keyword extraction.
Patent Information
- Application Number
- CN202510147262.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-07-04
AI Technical Summary
The existing text keyword extraction method lacks sufficient accuracy and representativeness when processing text with clear category labels, and fails to effectively improve the text keyword extraction accuracy under different category conditions.
By obtaining the content fields and category fields of pre-stored text, using antonym clustering and category field filtering, a keyword classification list corresponding to each category field is constructed, and the target text keywords under the specific category field are matched.
It improves the efficiency and accuracy of text keyword extraction under different categories, captures more potential text information with category colors, and solves the problem of insufficient text keyword extraction accuracy.
Smart Images

Figure CN120257981A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of natural language processing, and particularly to a method and system for extracting keywords of target text under a specific category. Background Art
[0002] In the era of information explosion, the rapid growth of text data has put forward higher requirements for information retrieval, content analysis, etc. As a core technology of text processing, text keyword extraction aims to automatically extract the most representative words or phrases from a large amount of text data, so as to facilitate users to quickly understand the text content, achieve text classification, abstract generation, or search optimization, etc.
[0003] Specifically, the existing text keyword extraction methods mainly include statistic-based extraction methods (such as TF-IDF), language model-based extraction methods (such as embedding models like Word2Vec, BERT, etc.), and graph model-based extraction methods (such as TextRank). These extraction methods can generally achieve good results, but when dealing with text with clear category labels, most of the above keyword extraction methods only rely on the statistical features or semantic features of the text itself, and often do not fully consider the specific category information to which the text belongs, which easily leads to the lack of sufficient accuracy and representativeness of the extracted text keywords.
[0004] Currently, no effective solution has been proposed for the problem of how to improve the accuracy of text keyword extraction under different category conditions in the related technology. Summary of the Invention
[0005] The embodiments of this application provide a method and system for extracting keywords of target text under a specific category, so as to at least solve the problem of how to improve the accuracy of text keyword extraction under different category conditions in the related technology.
[0006] In a first aspect, the embodiments of this application provide a method for extracting keywords of target text under a specific category, and the method includes:
[0007] Obtain pre-stored text, where the pre-stored text includes a content field and a category field;
[0008] Put the antonyms of the words in the rough keyword list of the pre-stored text into the word segmentation list of the pre-stored text, and cluster the words in the word segmentation list after putting them in to obtain a first cluster;
[0009] Based on the category field and the first cluster, screen the rough keyword list of the pre-stored text to obtain a candidate keyword list of the pre-stored text;
[0010] After obtaining the candidate keyword lists of multiple pre-stored texts, classify the keyword lists to obtain the keyword classification lists corresponding to each category field;
[0011] Based on the specific category field of the target text and the keyword classification lists corresponding to the respective category fields, match to obtain the keywords of the target text under the specific category field.
[0012] In some embodiments, put the antonyms of the words in the rough keyword list of the pre-stored text into the word segmentation list of the pre-stored text, and cluster the words in the word segmentation list after the putting to obtain the first clustering cluster including:
[0013] Perform word segmentation processing based on the content field of the pre-stored text, and preprocess the words obtained after the word segmentation processing to obtain the word segmentation list of the pre-stored text;
[0014] Put the antonyms of the words in the rough keyword list into the word segmentation list of the pre-stored text to obtain the word segmentation list after the putting, and cluster the words in the word segmentation list after the putting to obtain the first clustering cluster.
[0015] In some embodiments, based on the category field and the first clustering cluster, screen the rough keyword list of the pre-stored text to obtain the candidate keyword list of the pre-stored text including:
[0016] Extract keywords based on the content field of the pre-stored text to obtain the rough keyword list of the pre-stored text;
[0017] Based on the rough keyword list and the first clustering cluster, screen the rough keyword list to obtain the candidate keyword list of the pre-stored text.
[0018] In some embodiments, based on the rough keyword list and the first clustering cluster, screen the rough keyword list to obtain the candidate keyword list of the pre-stored text including:
[0019] Based on the rough keyword list and the first clustering cluster, determine whether the keywords in the rough keyword list and the antonyms of the keywords are in the same clustering cluster; and perform category color degree screening on the rough keyword list based on the result of the determination to obtain the initial candidate keyword list;
[0020] Based on the first clustering cluster, determine the outlier keywords in the initial candidate keyword list; and remove the outlier keywords from the candidate keyword list to obtain the candidate keyword list only containing non-outlier keywords;
[0021] Based on the non-outlier keywords in the candidate keyword list, the first clustering cluster is divided into a second clustering cluster and a third clustering cluster; and based on the third clustering cluster and the category field, category boundary screening is performed on the candidate keyword list that only contains non-outlier keywords to obtain a final candidate keyword list.
[0022] In some embodiments, based on the result of the judgment, category colorfulness screening is performed on the rough keyword list, and the initial candidate keyword list obtained includes:
[0023] If the keyword in the rough keyword list and the antonym of the keyword are in the same clustering cluster, it indicates that the keyword is a first keyword with unclear category colorfulness;
[0024] If the keyword in the rough keyword list and the antonym of the keyword are not in the same clustering cluster, it indicates that the keyword is a second keyword with obvious category colorfulness;
[0025] The first keyword in the rough keyword list is removed, and the second keyword is retained to obtain an initial candidate keyword list.
[0026] In some embodiments, based on the first clustering cluster, the outlier keywords in the initial candidate keyword list are determined to include:
[0027] Count the number of occurrences of the keywords in the initial candidate keyword list in the first clustering cluster. If the keyword only appears in one clustering cluster, the keyword is an outlier keyword.
[0028] In some embodiments, based on the non-outlier keywords in the candidate keyword list, dividing the first clustering cluster into a second clustering cluster and a third clustering cluster includes:
[0029] Based on the non-outlier keywords in the candidate keyword list, the in-cluster non-outlier words that are the same as the non-outlier keywords and different in-cluster common words are identified from the first clustering cluster;
[0030] The first clustering cluster that only contains in-cluster non-outlier words is denoted as the second clustering cluster, and the first clustering cluster that contains in-cluster common words and in-cluster non-outlier words is denoted as the third clustering cluster.
[0031] In some embodiments, based on the third clustering cluster and the category field, category boundary screening is performed on the candidate keyword list that only contains non-outlier keywords to obtain a final candidate keyword list, including:
[0032] Based on the category field, the in-cluster common words and the in-cluster non-outlier words in the third clustering cluster, calculate the first similarity between the category field and the in-cluster common words, and the second similarity between the category field and the in-cluster non-outlier words through a similarity calculation method;
[0033] Based on the first similarity and the second similarity, determine the in-cluster non-outlier words that do not meet the preset requirements, and remove the non-outlier keywords corresponding to the in-cluster non-outlier words from the candidate keyword list containing only non-outlier keywords to obtain the final candidate keyword list.
[0034] In some embodiments, after obtaining the candidate keyword lists of multiple pre-stored texts and performing a classification process on the keyword lists to obtain the keyword classification lists corresponding to each category field, the method further includes:
[0035] After obtaining the keyword classification lists corresponding to each category field, perform clustering screening and synonym expansion on the candidate keyword lists of the same category field respectively to obtain the updated keyword classification lists corresponding to each category field.
[0036] In some embodiments, based on the specific category field of the target text and the keyword classification lists corresponding to each category field, matching the keywords of the target text under the specific category field includes:
[0037] Based on the specific category field of the target text, confirm the keyword classification list under the specific category field from the keyword classification lists corresponding to each category field;
[0038] Based on the word segmentation result of the target text and the keyword classification list under the specific category field, match the keywords of the target text under the specific category field.
[0039] In a second aspect, an embodiment of the present application provides a target text keyword extraction system under a specific category, where the system is used to execute the method described in the first aspect above, and the system includes a classification list construction module and a keyword extraction module;
[0040] The classification list construction module is configured to obtain pre-stored text, where the pre-stored text includes a content field and a category field; put the antonyms of the words in the rough keyword list of the pre-stored text into the word segmentation list of the pre-stored text, and cluster the words in the word segmentation list after the putting to obtain a first cluster; based on the category field and the first cluster, screen the rough keyword list of the pre-stored text to obtain a candidate keyword list of the pre-stored text; after obtaining the candidate keyword lists of multiple pre-stored texts, perform a classification process on the keyword lists to obtain a keyword classification list corresponding to each category field.
[0041] The keyword extraction module is configured to match and obtain the keywords of the target text under the specific category field according to the specific category field of the target text and the keyword classification lists corresponding to the respective category fields.
[0042] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the method described in the first aspect above is implemented.
[0043] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in the first aspect above is implemented.
[0044] Compared with the related art, the present application provides a method and system for extracting target text keywords under a specific category. The method includes obtaining a pre-stored text containing content fields and category fields; putting the antonyms of the words in the rough keyword list of the pre-stored text into the word segmentation list of the pre-stored text, and clustering the words in the word segmentation list after putting them in to obtain the first clustering cluster; screening the rough keyword list of the pre-stored text based on the category field and the first clustering cluster to obtain a candidate keyword list of the pre-stored text; after obtaining the candidate keyword lists of multiple pre-stored texts, classifying the keyword lists to obtain a keyword classification list corresponding to each category field; and matching the specific category field of the target text with the keyword classification lists corresponding to each category field to obtain the keywords of the target text under the specific category field. This realizes constructing the keyword classification list corresponding to each category field based on the pre-stored text and matching the keywords of the target text under the specific category field, improving the efficiency of text keyword extraction under different category conditions. Moreover, in the construction of the keyword classification list corresponding to each category field, the introduction of antonym clustering to screen the rough keyword list is beneficial to capturing more potential text information with category characteristics, greatly improving the accuracy of text keyword extraction for the subsequent specific category field, and solving the problem of how to improve the accuracy of text keyword extraction under different category conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0046] Figure 1 is a flowchart of the steps of the method for extracting target text keywords under a specific category according to an embodiment of the present application;
[0047] Figure 2 is an internal structural schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be described and explained below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided by the present application without creative efforts belong to the scope of protection of the present application.
[0049] Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present application. For those of ordinary skill in the art, without creative efforts, the present application can also be applied to other similar scenarios based on these drawings. In addition, it can also be understood that although the efforts made in such a development process may be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing, or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be understood as the content disclosed in the present application being insufficient.
[0050] In the present application, the mention of "embodiment" means that the specific features, structures, or characteristics described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.
[0051] Unless otherwise defined, the technical terms or scientific terms involved in the present application should be the ordinary meanings understood by those with ordinary skills in the technical field to which the present application belongs. The words such as "a", "an", "one kind", "the", etc. involved in the present application do not represent a quantity limitation and can represent a single or plural number. The terms "including", "comprising", "having" and any variations thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device including a series of steps or modules (units) is not limited to the listed steps or units, but may further include unlisted steps or units, or may further include other steps or units inherent to these processes, methods, products, or devices. The terms "connected", "coupled", etc. involved in the present application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in the present application refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The terms "first", "second", "third", etc. involved in the present application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0052] The embodiment of the present application provides a method for extracting target text keywords in a specific category. Figure 1 It is a flowchart of the steps of the method for extracting target text keywords in a specific category according to the embodiment of the present application, asFigure 1 As shown in the figure, the method includes the following steps:
[0053] Step S102, obtain the pre-stored text, where the pre-stored text includes a content field and a category field;
[0054] Specifically for step S102, text data under various category conditions are collected to construct a text database. The text (i.e., the pre-stored text) in this database includes a content field and a category field. The text content of the text data obtained by collection is stored in this content field, and the corresponding text category is stored in this category field; step S102 directly obtains the pre-stored text from this text database.
[0055] In a preferred embodiment, the text targeted by the keyword extraction method of this preferred embodiment is medium and long text. In other words, for step S102, for example, news in text form under various categories (such as education, medical, real estate, finance, technology, entertainment, sports, games, automobiles, food, etc.) is obtained by using public news websites and web crawler technology, and it is judged whether the text length of the news is greater than or equal to a preset length, such as set to 100. If so, the text content of the news is stored in the content field of the text database, and the text category of the news is stored in the category field of the text database.
[0056] Step S104, put the antonyms of the words in the rough keyword list of the pre-stored text into the word segmentation list of the pre-stored text, and cluster the words in the word segmentation list after putting them in to obtain the first clustering cluster;
[0057] Step S104 specifically includes the following steps:
[0058] Step S1041, perform word segmentation processing based on the content field of the pre-stored text, and preprocess the words obtained after the word segmentation processing to obtain the word segmentation list of the pre-stored text;
[0059] In a preferred embodiment, step S1041 first initializes two empty lists, denoted as bg_ls_i and bg_ls_2_i respectively, where i represents the i-th pre-stored text; for the content field of the pre-stored text, a word segmentation extraction algorithm (such as jieba, jiagu, snownlp, hanlp and other algorithms) is used for word segmentation, and all the obtained words are stored in the list bg_ls_i. Further, download the Chinese stop word list, traverse each word in the list bg_ls_i. If the word is not in the Chinese stop word list, then store the word in the list bg_ls_2_i. After the traversal is completed, the words in the list bg_ls_2_i are de-duplicated, and the word segmentation list bg_ls_2_i of the pre-stored text is obtained after de-duplication.
[0060] Step S1042: Put the antonyms of the words in the rough keyword list into the word segmentation list of the pre-stored text to obtain the word segmentation list after putting, and cluster the words in the word segmentation list after putting to obtain the first clustering clusters.
[0061] In a preferred embodiment, step S1042 first initializes an empty list as bg_ls_3_i; downloads a Chinese antonym table, traverses each word in the rough keyword list simple_ls of the pre-stored text constructed in step S1061 below, and through this Chinese antonym table, finds any one antonym of each keyword in simple_ls (if there is no antonym, no processing is done. Since there are n keywords in simple_ls and there may be words without antonyms, the number of obtained antonyms is less than or equal to n), and combines the obtained antonyms with the words in the word segmentation list bg_ls_2_i of the pre-stored text and puts them into the list
[0062] bg_ls_3_i to obtain the word segmentation list bg_ls_3_i after putting. Further, use a clustering algorithm (such as
[0063] GMM, optics, dbscan, kmeans and other algorithms) to cluster the words in the word segmentation list bg_ls_3_i after putting to obtain t first clustering clusters.
[0064] Step S106: Based on the category field and the first clustering clusters, screen the rough keyword list of the pre-stored text to obtain the candidate keyword list of the pre-stored text;
[0065] Step S106 specifically includes the following steps:
[0066] Step S1061: Extract keywords based on the content field of the pre-stored text to obtain the rough keyword list of the pre-stored text;
[0067] In a preferred embodiment, for the content field of the pre-stored text in step S1061, use a word segmentation extraction algorithm (such as jieba, jiagu, snownlp, hanlp and other algorithms) to extract n keywords, and through 2 n-k-1 ≤tp_len_i≤2 n-k, k = tp_len_i / / p is used to calculate the value of n, where tp_len_i is the text length of the i-th pre-stored text, / / means taking the integer after division, and p is a preset positive integer. The parameter k is used to prevent the change of n from being insignificant as the text length of the pre-stored text increases. Further, an empty list simple_ls is constructed, and the obtained n keywords are sequentially placed into simple_ls. Finally, the updated list simple_ls is the rough keyword list simple_ls.
[0068] Regarding the parameter k, it should be noted by way of example that assuming k is not introduced, according to 2 n-1 ≤tp_len_i≤2 n , when tp_len_i is 300 and 500 respectively, the calculated n is 9 in both cases. Obviously, the increase in the text length of the pre-stored text does not cause the number of extracted keywords n to increase. In other words, by 2 n-k-1 ≤tp_len_i≤2 n-k to determine the number of extracted keywords, it can make the construction of the rough keyword list simple_ls more reasonable.
[0069] Step S1062: Based on the rough keyword list and the first clustering cluster, screen the rough keyword list to obtain the candidate keyword list of the pre-stored text.
[0070] Step S1062 specifically further includes the following steps:
[0071] Step S621: Based on the rough keyword list and the first clustering cluster, determine whether the keyword in the rough keyword list and the antonym of the keyword are in the same clustering cluster; and perform category colorfulness screening on the rough keyword list based on the judgment result to obtain the initial candidate keyword list;
[0072] In a preferred embodiment, the category colorfulness screening of the rough keyword list based on the judgment result to obtain the initial candidate keyword list is preferably:
[0073] If the keyword in the rough keyword list and the antonym of the keyword are in the same clustering cluster, it means the keyword is the first keyword with unclear category colorfulness; if the keyword in the rough keyword list and the antonym of the keyword are not in the same clustering cluster, it means the keyword is the second keyword with obvious category colorfulness; remove the first keyword in the rough keyword list and retain the second keyword to obtain the initial candidate keyword list.
[0074] It should be noted that an empty list is initialized as simple_ls_2. Each keyword in the rough keyword list simple_ls is traversed. Based on the t first clustering clusters obtained by the clustering algorithm in the above step S1042, each keyword will be in a corresponding clustering cluster. Further, if the keyword in simple_ls and its antonym are in the same clustering cluster, it reflects that the category color degree of the keyword itself is not obvious, so they are all distributed in the same clustering cluster and are recorded as the first keyword; if the keyword in simple_ls and its antonym are not in the same clustering cluster, it reflects that the category color degree of the keyword itself is obvious, so they are distributed in different clustering clusters and are recorded as the second keyword. Therefore, the second keyword is put into the empty list simple_ls_2. After the traversal is completed, the latest simple_ls_2 obtained is the candidate keyword list simple_ls_2.
[0075] Step S622: Based on the first clustering cluster, identify the outlier keywords in the initial candidate keyword list; and remove the outlier keywords from the candidate keyword list to obtain a candidate keyword list that only contains non-outlier keywords.
[0076] In a preferred embodiment, the outlier keywords in the initial candidate keyword list determined based on the first clustering cluster in step S622 are preferably:
[0077] Count the number of times the keywords in the initial candidate keyword list appear in the first clustering cluster. If a keyword appears in only one clustering cluster, then the keyword is an outlier keyword.
[0078] It should be noted that in step S622, a quantity distribution screening rule is introduced to screen the keywords in the candidate keyword list simple_ls_2 obtained in the above step S621. Specifically, based on the candidate keyword list simple_ls_2 in the above step S621, each keyword in simple_ls_2 is traversed, and it is successively verified whether each keyword also appears in the t first clustering clusters obtained in the above step S1042, and the number of clustering clusters where it appears is counted. If the statistical quantity is 1, it means that the keyword exists independently in a certain cluster and has an "outlier" attribute in the semantic information background. Therefore, the keyword is recorded as an outlier keyword, and the outlier keyword is removed from the initial candidate keyword list simple_ls_2 to obtain a candidate keyword list simple_ls_2 that only contains non-outlier keywords.
[0079] Step S623: Based on the non-outlier keywords in the candidate keyword list, the first clustering cluster is divided into a second clustering cluster and a third clustering cluster; and based on the third clustering cluster and the category field, a category boundary screening is performed on the candidate keyword list that only contains non-outlier keywords to obtain the final candidate keyword list.
[0080] In a preferred embodiment, in step S623, based on the non-outlier keywords in the candidate keyword list, dividing the first clustering cluster into a second clustering cluster and a third clustering cluster is preferably:
[0081] Based on the non-outlier keywords in the candidate keyword list, the in-cluster non-outlier words that are the same as the non-outlier keywords and the different in-cluster common words are identified from the first clustering cluster; the first clustering cluster that only contains in-cluster non-outlier words is denoted as the second clustering cluster, and the first clustering cluster that contains in-cluster common words and in-cluster non-outlier words is denoted as the third clustering cluster.
[0082] It should be noted that in step S623, a category boundary screening rule is introduced to screen the keywords in the candidate keyword list simple_ls_2 that only contains non-outlier keywords. Specifically, an empty list simple_ls_3_i is initialized, and each non-outlier keyword (denoted as hk_word_i) in the candidate keyword list simple_ls_2 that only contains non-outlier keywords is traversed. There are two situations for the distribution of the first clustering cluster where the hk_word_i is located: ① If all the words in the cluster are non-outlier keywords (in-cluster non-outlier words), it means that the first clustering cluster has a precise and strong category color in the semantic information background of the target text. The first clustering cluster is denoted as the second clustering cluster, and the hk_word_i is stored in simple_ls_3_i; ② If there are some words in the first clustering cluster that are not non-outlier keywords (in-cluster common words), the first clustering cluster is denoted as the third clustering cluster.
[0083] In a preferred embodiment, in step S623, based on the third clustering cluster and the category field, performing a category boundary screening on the candidate keyword list that only contains non-outlier keywords to obtain the final candidate keyword list is preferably:
[0084] Based on the category field, the in-cluster common words and in-cluster non-outlier words in the third clustering cluster, through a similarity calculation method, calculate the first similarity between the category field and the in-cluster common words, and the second similarity between the category field and the in-cluster non-outlier words; based on the first similarity and the second similarity, determine the in-cluster non-outlier words that do not meet the preset requirements, and remove the non-outlier keywords corresponding to the in-cluster non-outlier words from the candidate keyword list that only contains non-outlier keywords to obtain the final candidate keyword list.
[0085] It should be noted that in step S623, a category boundary screening rule is introduced to screen the keywords in the candidate keyword list simple_ls_2 that only contain non-outlier keywords. A word vector conversion model (such as fasttext, gensim, roberta, etc.) is introduced to convert the category field of the pre-stored text into a p-dimensional word vector (denoted as class_vect), and the in-cluster non-outlier words (hk_word_i) in the third clustering cluster are also converted into p-dimensional word vectors (denoted as hk_vect_i). The in-cluster common words in the third clustering cluster are also converted into p-dimensional word vectors, and the Euclidean distance between class_vect and the p-dimensional word vectors of the in-cluster common words is calculated (it is also possible to calculate similarity metric coefficients such as the cosine distance, Manhattan distance, Chebyshev distance, Hamming distance, and Mahalanobis distance between class_vect and the p-dimensional word vectors of the in-cluster common words, which will not be elaborated here one by one). The Euclidean distance calculation formula is as follows:
[0086]
[0087] where f i represents the value of the i-th dimension in class_vect, and y i represents the value of the i-th dimension in the p-dimensional word vector of the in-cluster common words. If there are m in-cluster common words under this third clustering cluster, m Euclidean distances can be obtained according to this Euclidean distance calculation formula, and the minimum value among the m Euclidean distances is taken, denoted as flag_D.
[0088] Calculate the Euclidean distance between class_vect and the p-dimensional word vectors of the in-cluster non-outlier words. The Euclidean distance calculation formula is as follows:
[0089]
[0090] where f i represents the value of the i-th dimension in class_vect, and z i represents the value of the i-th dimension in the p-dimensional word vector of the in-cluster non-outlier words.
[0091] Further, compare the D_z of the non-outlier words within the cluster with the above flag_D. If D_z ≥ flag_D, it indicates that the similarity between the non-outlier words within the cluster and the category to which the pre-stored text belongs does not exceed that of all ordinary words within the cluster. Therefore, it is determined that the non-outlier words within the cluster do not meet the preset requirements for becoming text keywords under this category field, and the non-outlier words within the cluster are not stored in simple_ls_3_i (or directly removed from the above candidate keyword list simple_ls_2 that only contains non-outlier keywords); if D_z < flag_D, it indicates that the similarity between the non-outlier words within the cluster and the category to which the pre-stored text belongs exceeds that of all ordinary words within the cluster. Therefore, it is determined that the non-outlier words within the cluster meet the preset requirements for becoming text keywords under this category field, and the non-outlier words within the cluster are stored in simple_ls_3_i. After the traversal is completed, the final candidate keyword list simple_ls_3_i is obtained.
[0092] Step S108, after obtaining the candidate keyword lists of multiple pre-stored texts, classify the keyword lists to obtain the keyword classification lists corresponding to each category field;
[0093] In a preferred embodiment, after obtaining the candidate keyword list simple_ls_3_i of multiple pre-stored texts in step S108, the number of categories class_j of the multiple pre-stored texts is counted. For example, assume that the counted number of class_j is 10 (such as calss_1 = education, class_2 = medical, class_3 = real estate, class_4 = finance and economics, class_5 = technology, class_6 = entertainment, class_7 = sports, class_8 = games, class_9 = automobiles, class_10 = food), then the category of each pre-stored text must belong to one of these 10 categories; initialize 10 empty keyword classification lists key_cluster_list_j, obtain all the candidate keyword lists simple_ls_3_i under each class_j, extract the words in the obtained all candidate keyword lists simple_ls_3_i in sequence and store them into key_cluster_list_j, and remove duplicates of the repeated words in key_cluster_list_j to obtain the keyword classification list key_cluster_list_j. Finally, the keyword classification lists corresponding to each category field are obtained, that is, key_cluster_list_1, key_cluster_list_2, key_cluster_list_3, key_cluster_list_4, key_cluster_list_5, key_cluster_list_6, key_cluster_list_7, key_cluster_list_8, key_cluster_list_9, and key_cluster_list_10 are obtained respectively.
[0094] Optionally, after obtaining the keyword classification lists corresponding to each category field, clustering and screening and synonym expansion are respectively performed on the candidate keyword lists of the same category field to obtain the updated keyword classification lists corresponding to each category field.
[0095] In a preferred embodiment, a clustering algorithm (such as GMM, optics, dbscan, kmeans, etc.) is used to cluster all the keywords in the keyword classification list key_cluster_list_j corresponding to each of the above-obtained category fields in sequence, and h_j corresponding clusters are obtained in sequence; for the h_j clusters obtained, the number of keywords in each cluster is counted to obtain the median (for example, for the current key_cluster_list_5, h_5 = 7, and the numbers of keywords in the 7 clusters are 22, 10, 8, 50, 100, 200, 150 respectively, and the median of this cluster is statistically obtained as 50). Further, if the number of keywords in the cluster is greater than the median, it indicates that the keywords in the cluster can show relatively strong and accurate category semantic information for the text, and all the keywords in the cluster are retained in key_cluster_list_j; if the number of keywords in the cluster is less than or equal to the median, it indicates that the keywords in the cluster cannot show relatively strong and accurate category semantic information for the text, and all the keywords in the cluster are removed from key_cluster_list_j.
[0096] For the key_cluster_list_j obtained by median filtering after the above clustering, keyword expansion is performed. Specifically, a Chinese synonym table is downloaded. For each key_cluster_list_j, each keyword in the key_cluster_list_j is traversed, and all synonyms corresponding to each keyword are obtained according to the Chinese synonym table; and all synonyms of each keyword are stored in the key_cluster_list_j. By introducing synonyms based on the Chinese synonym table, a certain number of keywords are generated to achieve the effect of keyword number expansion, which can make up for the problem that the number of keywords in key_cluster_list_j tends to be small in some cases.
[0097] Step S110: Based on the specific category field of the target text and the keyword classification lists corresponding to each category field, the keywords of the target text under the specific category field are matched.
[0098] Specifically, in step S110, based on the specific category field of the target text, the keyword classification list under the specific category field is confirmed from the keyword classification lists corresponding to each category field; based on the word segmentation result of the target text and the keyword classification list under the specific category field, the keywords of the target text under the specific category field are matched.
[0099] In a preferred embodiment, in step S110, the target text (assumed to be goal_text) is obtained. If it is desired to obtain several keywords of goal_text under a specific category field, first, the specific category field is determined. Based on this specific category field, from the keyword classification lists key_cluster_list_j corresponding to each category field, the keyword classification list under this specific category field is matched. Secondly, an empty list goal_list is initialized for the target text goal_text. The goal_text is segmented by a word segmentation extraction algorithm (such as jieba, snownlp, jiagu, etc.), and each obtained word ci_i and its corresponding importance imp_i are stored in the same tuple goal_i = (ci_i, imp_i). Further, all goal_i are stored in goal_list, and the order of goal_i in goal_list is adjusted according to the rule of arranging from largest to smallest imp_i. Thirdly, each goal_i in goal_list is traversed, and ci_i is taken out. If ci_i exists in key_cluster_list_j, then goal_i is retained in goal_list. If ci_i does not exist in key_cluster_list_j, then goal_i is removed from goal_list. Finally, the ci_i in the obtained goal_list are the keywords of the target text goal_text under the specific category field, and at the same time, the sorting from largest to smallest based on importance is achieved.
[0100] Through the above steps in the embodiments of the present application, it is realized to first construct the keyword classification lists corresponding to each category field based on the pre-stored text, match the keywords of the target text under the specific category field, improve the text keyword extraction efficiency under different category conditions, and in the construction of the keyword classification lists corresponding to each category field, the antonym clustering is introduced to screen the rough keyword list, which is beneficial to capturing more potential text information with category colors, greatly improving the accuracy of the text keyword extraction of the subsequent specific category field, and solving the problem of how to improve the accuracy of the text keyword extraction under different category conditions.
[0101] It should be noted that the steps shown in the above process or the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0102] The embodiments of the present application provide a target text keyword extraction system under a specific category. The system includes a classification list construction module and a keyword extraction module;
[0103] A classification list construction module is configured to obtain pre-stored texts, where the pre-stored texts include content fields and category fields; put the antonyms of the words in the rough keyword list of the pre-stored texts into the word segmentation list of the pre-stored texts, and cluster the words in the word segmentation list after putting them in to obtain a first cluster; based on the category fields and the first cluster, screen the rough keyword list of the pre-stored texts to obtain a candidate keyword list of the pre-stored texts; after obtaining the candidate keyword lists of multiple pre-stored texts, classify the keyword lists to obtain a keyword classification list corresponding to each category field.
[0104] A keyword extraction module is configured to match and obtain the keywords of the target text under a specific category field according to the specific category field of the target text and the keyword classification lists corresponding to each category field.
[0105] Through the classification list construction module and the keyword extraction module in the embodiments of the present application, it is realized to first construct a keyword classification list corresponding to each category field based on the pre-stored texts, match and obtain the keywords of the target text under a specific category field, improve the text keyword extraction efficiency under different category conditions, and in the construction of the keyword classification lists corresponding to each category field, the antonym clustering is introduced to screen the rough keyword list, which is beneficial to capturing more potential text information with category colors, greatly improving the accuracy of the text keyword extraction of the subsequent specific category fields, and solving the problem of how to improve the accuracy of the text keyword extraction under different category conditions.
[0106] It should be noted that the above-mentioned each module can be a functional module or a program module, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned each module can be located in the same processor; or the above-mentioned each module can also be located in different processors in any combination form.
[0107] This embodiment also provides an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0108] Optionally, the above-mentioned electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above-mentioned processor, and the input / output device is connected to the above-mentioned processor.
[0109] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be repeated here.
[0110] In addition, in combination with the target text keyword extraction method under a specific category in the above embodiments, an embodiment of the present application can provide a storage medium to implement this. A computer program is stored on the storage medium; when the computer program is executed by a processor, the target text keyword extraction method under any one of the above embodiments is implemented.
[0111] In one embodiment, a computer device is provided. The computer device may be a terminal. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a target text keyword extraction method under a specific category is implemented. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0112] In one embodiment, Figure 2 is a schematic internal structure diagram of an electronic device according to an embodiment of the present application. As Figure 2 shown, an electronic device is provided. The electronic device may be a server, and its internal structure diagram may be as Figure 2 shown. The electronic device includes a processor, a network interface, an internal memory, and a non-volatile memory connected through an internal bus. Among them, the non-volatile memory stores an operating system, a computer program, and a database. The processor is used to provide computing and control capabilities. The network interface is used to communicate with an external terminal through a network connection. The internal memory is used to provide an environment for the operation of the operating system and the computer program. When the computer program is executed by the processor, a target text keyword extraction method under a specific category is implemented. The database is used to store data.
[0113] Those skilled in the art can understand that Figure 2 the structure shown in
[0114] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0115] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0116] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A method for extracting target text keywords under a specific category, characterized in that, The method includes: Obtain a pre-stored text, where the pre-stored text includes a content field and a category field; Put the antonyms of the words in the rough keyword list of the pre-stored text into the word segmentation list of the pre-stored text, and cluster the words in the word segmentation list after putting, to obtain a first cluster; Based on the category field and the first cluster, screen the rough keyword list of the pre-stored text to obtain a candidate keyword list of the pre-stored text; After obtaining the candidate keyword lists of multiple pre-stored texts, classify the keyword lists to obtain keyword classification lists corresponding to each category field; Based on the specific category field of the target text and the keyword classification lists corresponding to each category field, match to obtain the keywords of the target text under the specific category field.
2. The method according to claim 1, characterized in that, Put the antonyms of the words in the rough keyword list of the pre-stored text into the word segmentation list of the pre-stored text, and cluster the words in the word segmentation list after putting, to obtain a first cluster, including: Perform word segmentation processing based on the content field of the pre-stored text, and preprocess the words obtained after the word segmentation processing to obtain the word segmentation list of the pre-stored text; Put the antonyms of the words in the rough keyword list into the word segmentation list of the pre-stored text to obtain the word segmentation list after putting, and cluster the words in the word segmentation list after putting to obtain a first cluster.
3. The method according to claim 1, characterized in that, Based on the category field and the first cluster, screen the rough keyword list of the pre-stored text to obtain a candidate keyword list of the pre-stored text, including: Extract keywords based on the content field of the pre-stored text to obtain a rough keyword list of the pre-stored text; Based on the rough keyword list and the first cluster, screen the rough keyword list to obtain a candidate keyword list of the pre-stored text.
4. The method according to claim 3, wherein Based on the rough keyword list and the first cluster, screen the rough keyword list to obtain a candidate keyword list of the pre-stored text, including: Based on the rough keyword list and the first cluster, determine whether the keyword in the rough keyword list and the antonym of the keyword are in the same cluster; and based on the result of the determination, perform category color degree screening on the rough keyword list to obtain an initial candidate keyword list; Based on the first cluster, determine the outlier keywords in the initial candidate keyword list; and remove the outlier keywords from the candidate keyword list to obtain a candidate keyword list only containing non-outlier keywords; Based on the non-outlier keywords in the candidate keyword list, divide the first cluster into a second cluster and a third cluster; and based on the third cluster and the category field, perform category boundary screening on the candidate keyword list only containing non-outlier keywords to obtain a final candidate keyword list.
5. The method according to claim 4, wherein Perform category colorfulness screening on the rough keyword list based on the result of the judgment, and obtain an initial candidate keyword list including: If the keyword in the rough keyword list and the antonym of the keyword are in the same clustering cluster, it indicates that the keyword is a first keyword with unclear category colorfulness; If the keyword in the rough keyword list and the antonym of the keyword are not in the same clustering cluster, it indicates that the keyword is a second keyword with obvious category colorfulness; Remove the first keywords in the rough keyword list and retain the second keywords to obtain an initial candidate keyword list.
6. The method according to claim 4, wherein Based on the first clustering cluster, determine the outlier keywords in the initial candidate keyword list including: Count the number of occurrences of the keywords in the initial candidate keyword list in the first clustering cluster. If the keyword only appears in one clustering cluster, the keyword is an outlier keyword.
7. The method according to claim 4, characterized in that Based on the non-outlier keywords in the candidate keyword list, divide the first clustering cluster into a second clustering cluster and a third clustering cluster including: Based on the non-outlier keywords in the candidate keyword list, identify the in-cluster non-outlier words that are the same as the non-outlier keywords and the in-cluster common words that are different from the non-outlier keywords in the first clustering cluster; Record the first clustering cluster that only contains in-cluster non-outlier words as the second clustering cluster, and record the first clustering cluster that contains in-cluster common words and in-cluster non-outlier words as the third clustering cluster.
8. The method according to claim 7, wherein Based on the third clustering cluster and the category field, perform category boundary screening on the candidate keyword list that only contains non-outlier keywords, and obtain a final candidate keyword list including: Based on the category field, the in-cluster common words and in-cluster non-outlier words in the third clustering cluster, calculate the first similarity between the category field and the in-cluster common words, and the second similarity between the category field and the in-cluster non-outlier words through a similarity calculation method; Based on the first similarity and the second similarity, determine the in-cluster non-outlier words that do not meet the preset requirements, and remove the non-outlier keywords corresponding to the in-cluster non-outlier words from the candidate keyword list that only contains non-outlier keywords to obtain a final candidate keyword list.
9. The method according to claim 1, characterized in that After obtaining the candidate keyword lists of multiple pre-stored texts, and performing classification processing on the keyword lists to obtain the keyword classification lists corresponding to each category field, the method further includes: After obtaining the keyword classification lists corresponding to each category field, perform clustering screening and synonym expansion on the candidate keyword lists of the same category field respectively to obtain the updated keyword classification lists corresponding to each category field.
10. The method according to claim 1, wherein Based on the specific category field of the target text and the keyword classification lists corresponding to each category field, match the keywords of the target text under the specific category field including: Based on the specific category field of the target text, confirm the keyword classification list under the specific category field from the keyword classification lists corresponding to each category field; Based on the word segmentation result of the target text and the keyword classification list under the specific category field, the keywords of the target text under the specific category field are obtained by matching.
11. A target text keyword extraction system under a specific category, characterized in that, The system is used to execute the method according to any one of claims 1 to 10. The system includes a classification list construction module and a keyword extraction module. The classification list construction module is used to obtain the pre-stored text, where the pre-stored text includes a content field and a category field; put the antonyms of the words in the rough keyword list of the pre-stored text into the word segmentation list of the pre-stored text, and cluster the words in the word segmentation list after putting to obtain the first clustering cluster; based on the category field and the first clustering cluster, screen the rough keyword list of the pre-stored text to obtain the candidate keyword list of the pre-stored text; after obtaining the candidate keyword lists of multiple pre-stored texts, classify the keyword lists to obtain the keyword classification lists corresponding to each category field. The keyword extraction module is used to obtain the keywords of the target text under the specific category field by matching according to the specific category field of the target text and the keyword classification lists corresponding to each category field.
12. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is set to run the computer program to execute the method according to any one of claims 1 to 10.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1 to 10.