Automatic keyword recommendation method and system based on semantics
Through data processing, word segmentation processing, keyword processing and encoding, combined with Chinese word segmentation and word embedding models, the problem of narrow keyword recommendation results in the existing technology is solved, and efficient and accurate automatic keyword recommendation is achieved to adapt to corpus of different fields and scales.
Patent Information
- Application Number
- CN202510378627.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-18
AI Technical Summary
When existing keyword recommendation methods deal with complex or ambiguity recommendation matches, the generated recommendation results are too narrow to cover the diversity of recommendation intentions, and it is difficult to achieve efficient and accurate automatic keyword recommendations.
Through data processing, word segmentation processing, keyword processing and keyword encoding, combined with Chinese word segmentation, word embedding model and vector database, keywords are screened using bipartite search method and primary and secondary factor analysis method, and recommended through vector similarity search.
It realizes efficient and accurate extraction of keyword information in large-scale corpus, adapts to corpus of different fields and scales, is universal and extensible, and ensures the timeliness and diversity of data during the periodic maintenance process.
Smart Images

Figure CN120336549A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and particularly to a method and system for automatically recommending keywords based on semantics. Background Art
[0002] Keyword recommendation can automatically recommend or associate potential interesting and more product- or service-content-relevant similar tags according to the keyword tags given by users, so as to ensure that users are more likely to browse content more relevant to their interests, thereby improving the visibility of content to users.
[0003] Existing keyword recommendation methods rely to a large extent on keyword hit or fuzzy query methods. Although this can solve relatively simple recommendation tasks, due to complete reliance on strict exact matching or only relying on surface semantics, when dealing with complex or ambiguous recommendation matches, the generated recommendation results are too narrow to cover the diversity of recommendation intents and it is difficult to achieve ideal implementation results.
[0004] How to efficiently and accurately implement the task of automatic keyword recommendation is a technical problem to be solved. Summary of the Invention
[0005] The technical task of the present invention is to, aiming at the above deficiencies, provide a method and system for automatically recommending keywords based on semantics to solve the technical problem of how to efficiently and accurately implement the task of automatic keyword recommendation.
[0006] In a first aspect, a method for automatically recommending keywords based on semantics according to the present invention includes the following steps:
[0007] Data processing: Collect text data according to task requirements, store the text data in a first corpus, perform data cleaning on the text data to be cleaned in the first corpus, and supplement the cleaned text data to a second corpus;
[0008] Word segmentation processing: Perform word segmentation processing on the text data in the second corpus, update the statistical data of the words recorded in the first word segmentation dataset according to the words obtained by the word segmentation processing, supplement unrecorded words and the corresponding part-of-speech and statistical data of the unrecorded words to the first word segmentation dataset, screen the words recorded in the first word segmentation dataset, and calculate the TF-IDF values of the recorded words in combination with the scale of the text data in the second corpus, and record the recorded words and the corresponding TF-IDF values in a second word segmentation dataset, where, for the recorded words and unrecorded words, the statistical data includes the number of occurrences of the word and the number of texts containing the word;
[0009] Keyword processing: Use binary search and primary and secondary factor analysis to select available keywords from the second segmentation data set, and add the selected available keywords to the keyword word list;
[0010] Keyword encoding: Encode the keywords in the keyword vocabulary through the word embedding model to obtain the corresponding encoding vector, and add the keywords and the corresponding encoding vector to the keyword set of the vector database;
[0011] Keyword recommendation: For a given word, encode the given word through the word embedding model to obtain the encoding vector corresponding to the given word. Based on the encoding vector corresponding to the given word, retrieve keywords with similarity matching from the keyword set of the vector database through vector similarity search. For the retrieved words, delete the words that are exactly the same as the given words according to task requirements.
[0012] Preferably, the data processing comprises the following steps:
[0013] According to business needs, determine the type and scope of text data to be collected, collect text data based on the determined type and scope, and aggregate the collected text data into a first forecast library;
[0014] The text data to be cleaned in the first corpus is read regularly by a program script, a data cleaning operation is performed on the text data to be cleaned, and the cleaned text data is added to the second corpus, wherein the data cleaning operation includes removing duplicate data, removing irrelevant characters, and removing noise data.
[0015] Preferably, the word segmentation process includes the following steps:
[0016] The text data to be segmented in the second corpus is read in sequence, the Jieba segmentation algorithm is called through a program script, the read text data is segmented when the segmentation mode is the search mode, and stop words are removed in combination with a preset stop word list during the segmentation process, and there are two types of words obtained by the segmentation, namely, words recorded in the first segmentation data set and words not recorded in the first segmentation data set, for the words recorded in the first segmentation data set, the statistical data corresponding to the recorded words in the first segmentation data set are updated, and for the words not recorded in the segmentation data set, the unrecorded words and the part of speech and statistical data corresponding to the unrecorded words are added to the first segmentation data set;
[0017] For the first word segmentation dataset, all recorded words in the first word segmentation dataset are screened as a whole according to the set maintenance cycle and based on the set target part-of-speech range, and statistical data corresponding to the screened recorded words are read in sequence, and the TF-IDF of the read recorded words is calculated in combination with the scale information of the text data in the second corpus, and the recorded words and their corresponding TF-IDF values are recorded in the second word segmentation dataset;
[0018] Among them, the calculation formula for the TF-IDF value of the recorded words is as follows:
[0019]
[0020] Preferably, the keyword processing includes the following steps:
[0021] If the keyword set has not been created, construct the keyword set;
[0022] Sort the recorded words in the second word segmentation data set in descending order according to the TF-IDF value;
[0023] Based on the sorting result, use the ordered integer set composed of the sorting serial numbers as the initial search set of the binary search algorithm, obtain the TF-IDF values of the recorded words whose serial numbers are less than or equal to the specified truncation value, calculate the proportion of the sum of the obtained TF-IDF values of the recorded words, use the proportion as the cumulative frequency index applied in the primary and secondary factor analysis method, combine the binary search algorithm and the primary and secondary factor analysis method to determine the available keywords in the second word segmentation data set, and supplement the unrecorded available keywords to the keyword list;
[0024] Among them, the calculation formula for the proportion is:
[0025]
[0026] In a second aspect, an automatic keyword recommendation system based on semantics according to the present invention is used to realize automatic keyword recommendation through an automatic keyword recommendation method according to any one of the first aspect. The system includes a data processing module, a word segmentation processing module, a keyword processing module, a keyword encoding module, and a keyword recommendation module;
[0027] The data processing module is used to perform the following: collect text data according to the task requirements, store the text data in the first corpus, perform data cleaning on the text data to be cleaned in the first corpus, and supplement the cleaned text data to the second corpus;
[0028] The word segmentation processing module is used to perform the following: perform word segmentation on the text data in the second corpus, update the statistical data corresponding to the words recorded in the first word segmentation dataset according to the words obtained from the word segmentation, supplement the first word segmentation dataset with unrecorded words and the corresponding part-of-speech and statistical data of the unrecorded words, screen the words recorded in the first word segmentation dataset, and calculate the TF-IDF values of the recorded words in combination with the scale of the text data in the second corpus, and record the recorded words and the corresponding TF-IDF values in the second word segmentation dataset. Among them, for the recorded words and unrecorded words, the statistical data includes the number of occurrences of the word and the number of texts containing the word;
[0029] The keyword processing module is used to perform the following: screen available keywords from the second word segmentation dataset through the binary search method and the primary and secondary factor analysis method, and supplement the selected available keywords to the keyword list;
[0030] The keyword encoding module is used to perform the following: encode the keywords in the keyword list through a word embedding model to obtain the corresponding encoding vectors, and supplement the keywords and the corresponding encoding vectors to the keyword set of the vector database;
[0031] The keyword recommendation module is used to perform the following: for a given word, encode the given word through a word embedding model to obtain the encoding vector corresponding to the given word, retrieve the keywords with similarity matching from the keyword set of the vector database based on the encoding vector corresponding to the given word through vector similarity search, and for the retrieved words, delete the words that are exactly the same as the given word according to the task requirements.
[0032] Preferably, the data processing module is used to perform the following operations:
[0033] According to the business requirements, determine the type and scope of the text data to be collected, collect the text data based on the determined type and scope, and gather the collected text data into the first corpus;
[0034] Regularly read the text data to be cleaned in the first corpus through a program script, perform data cleaning operations on the text data to be cleaned, and supplement the cleaned text data to the second corpus. Among them, the data cleaning operations include removing duplicate data, removing irrelevant characters, and removing noise data.
[0035] Preferably, the word segmentation processing module is used to perform the following operations:
[0036] Read the text data to be segmented in the second corpus in sequence, call the Jieba segmentation algorithm through a program script, segment the read text data in the case of the segmentation mode being the search mode, and remove stop words in combination with a preset stop word list during the segmentation process. There are two types of words obtained by segmentation, namely the words already recorded in the first segmentation dataset and the words not recorded in the first segmentation dataset. For the words already recorded in the first segmentation dataset, update the statistical data corresponding to the words already recorded in the first segmentation dataset. For the words not recorded in the segmentation dataset, supplement the unrecorded words and their corresponding parts of speech and statistical data to the first segmentation dataset;
[0037] For the first segmentation dataset, according to the set maintenance period and based on the set target part-of-speech range, screen all the words already recorded in the first segmentation dataset as a whole, and read the statistical data corresponding to the screened words already recorded in sequence. Combine the scale information of the text data in the second corpus to calculate the TF-IDF of the read words already recorded, and record the already recorded words and their corresponding TF-IDF values in the second segmentation dataset;
[0038] Among them, the calculation formula for the TF-IDF value of the words already recorded is as follows:
[0039]
[0040] Preferably, the keyword processing module is used to perform the following:
[0041] If the keyword set has not been created, construct the keyword set;
[0042] Sort the words already recorded in the second segmentation dataset in descending order according to the TF-IDF value;
[0043] Based on the sorting result, use the ordered integer set composed of the sorting serial numbers as the initial search set of the binary search algorithm, obtain the TF-IDF values of the words already recorded with the serial number less than or equal to the specified truncation value, calculate the proportion of the sum of the obtained TF-IDF values of the words already recorded, use the proportion as the cumulative frequency index applied in the primary and secondary factor analysis method, combine the binary search algorithm and the primary and secondary factor analysis method to determine the available keywords in the second segmentation dataset, and supplement the unrecorded available keywords to the keyword table;
[0044] Among them, the calculation formula for the proportion is:
[0045]
[0046] The semantic-based keyword automatic recommendation method and system of the present invention have the following advantages:
[0047] 1. The combination of Chinese word segmentation and keyword generation technology can efficiently extract keyword information from a large-scale target corpus on the basis of accurately capturing the topic information of the corpus.
[0048] 2. Based on the word embedding model and vector database application technology, it can achieve efficient and highly relevant matching recommendations on the basis of deeply understanding the semantics of keywords.
[0049] 3. It can adapt to corpora in different fields and scales, has strong generality and scalability, and the method synchronously realizes the periodic maintenance process of the recommendation task, which can ensure the timeliness of data and increase the diversity of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0051] The present invention will be further described below with reference to the drawings.
[0052] Figure 1 FIG. is a flowchart of a semantic-based automatic keyword recommendation method for Embodiment 1;
[0053] Figure 2 FIG. is a flowchart of the processing of Chinese part data of a semantic-based automatic keyword recommendation method for Embodiment 1;
[0054] Figure 3 FIG. is a full flowchart of automatic keyword recommendation of a semantic-based automatic keyword recommendation method for Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] The present invention will be further described below with reference to the drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it. However, the specific embodiments cited do not limit the present invention. Without conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0056] The embodiments of the present invention provide a semantic-based automatic keyword recommendation method and system, which are used to solve the technical problem of how to efficiently and accurately implement the automatic keyword recommendation task.
[0057] Embodiment 1:
[0058] A semantic-based automatic keyword recommendation method of the present invention includes five steps: data processing, word segmentation processing, keyword processing, keyword encoding, and keyword recommendation.
[0059] Step S100 Data Processing: Collect text data according to task requirements, store the text data in the first corpus, perform data cleaning on the text data to be cleaned in the first corpus, and supplement the cleaned text data to the second corpus.
[0060] As a specific implementation of data processing, this step includes the following operations:
[0061] (1) According to business requirements, determine the type and scope of the text data to be collected, collect the text data based on the determined type and scope, and gather the collected text data into the first corpus;
[0062] (2) Regularly read the text data to be cleaned in the first corpus through a program script, perform data cleaning operations on the text data to be cleaned, and supplement the cleaned text data to the second corpus. Among them, the data cleaning operations include removing duplicate data, removing irrelevant characters, and removing noise data.
[0063] When this step is executed, if the first corpus and the second corpus have not been constructed, complete the initial construction. According to the specific task requirements, determine the detailed types and collection scopes of the text data to be collected, and accordingly complete the collection of text data, which is summarized into the first corpus. Read the data to be cleaned in the first corpus through a program script for a series of data cleaning processes, and supplement the obtained text data to the second corpus. The data cleaning process for the text data includes removing duplicate text data, removing irrelevant characters, and removing noise data, etc.
[0064] As Figure 3 shown, for the initial implementation process of the task, based on the initially collected text data, the first corpus and the second corpus can be directly constructed according to the method described in this step; for the subsequent periodic maintenance process of the task, according to the set maintenance period, complete the supplement and maintenance of the second corpus based on the newly added data to be cleaned in the first corpus according to the method described in this step.
[0065] Step S200 Word Segmentation Processing: Perform word segmentation on the text data in the second corpus, update the statistical data corresponding to the words recorded in the first word segmentation dataset according to the words obtained from the word segmentation, and supplement the unrecorded words and the corresponding part-of-speech and statistical data of the unrecorded words to the first word segmentation dataset. Screen the words recorded in the first word segmentation dataset, and combine the scale of the text data in the second corpus to calculate the TF-IDF values of the recorded words, and record the recorded words and the corresponding TF-IDF values in the second word segmentation dataset. Among them, for the recorded words and the unrecorded words, the statistical data includes the number of occurrences of the word and the number of texts containing the word.
[0066] As a specific implementation of word segmentation processing, this step includes the following operations:
[0067] (1) Read the text data to be segmented in the second corpus in sequence. Call the Jieba segmentation algorithm through a program script, segment the read text data in the case of the segmentation mode being the search mode, and remove stop words in combination with a preset stop word list during the segmentation process. There are two types of words obtained by segmentation, namely the words already recorded in the first segmentation dataset and the words not recorded in the first segmentation dataset. For the words already recorded in the first segmentation dataset, update the statistical data corresponding to the words already recorded in the first segmentation dataset. For the words not recorded in the segmentation dataset, supplement the unrecorded words, the corresponding part-of-speech, and the statistical data to the first segmentation dataset;
[0068] (2) For the first segmentation dataset, overall screen all the words already recorded in the first segmentation dataset according to the set maintenance period and based on the set target part-of-speech range, and read the statistical data corresponding to the screened words already recorded in sequence. Combine the scale information of the text data in the second corpus, calculate the TF-IDF of the read words already recorded, and record the already recorded words and their corresponding TF-IDF values in the second segmentation dataset;
[0069] Among them, the calculation formula for the TF-IDF value of the words already recorded is as follows:
[0070]
[0071] In this step, segment the data to be segmented in the second corpus, determine the first segmentation dataset, and obtain the second segmentation dataset through target part-of-speech screening and in combination with the improved TF-IDF method. If the first segmentation dataset has not been constructed yet, complete the initial construction. Read the data to be segmented in the second corpus in sequence, use the Chinese word segmentation method to segment the read text data, and then based on the words obtained by segmentation, update the statistical data corresponding to the words already recorded in the first segmentation dataset, and supplement the unrecorded words, the corresponding part-of-speech, and their corresponding statistical data to the first segmentation dataset; for the first segmentation dataset, overall screen all the words already recorded in this dataset according to the set target part-of-speech range, and read the statistical data corresponding to the screened words already recorded in sequence, and combine the scale information of the text data in the second corpus, calculate the TF-IDF of the read words already recorded according to the corresponding formula, and record the already recorded words and their corresponding TF-IDF values in the second segmentation dataset.
[0072] Among them, the process of text data segmentation includes: calling the Jieba segmentation algorithm through a program script, segmenting the read text data in the case of the segmentation mode being the search mode, and removing stop words in combination with a preset stop word list during the segmentation process to obtain the final segmented data.
[0073] Specifically, such as Figure 3As shown, for the initial implementation process of the task, the first word segmentation dataset and the second word segmentation dataset can be determined in sequence according to the method described in this step; for the subsequent cycle maintenance process of the task, according to the set maintenance cycle, the newly supplemented data to be word-segmented in the second corpus is used to complete the supplementation and maintenance of the first word segmentation dataset, and the second word segmentation dataset is reconstructed accordingly.
[0074] Step S300 Keyword processing: Screen available keywords from the second word segmentation dataset through the binary search method and the primary and secondary factor analysis method, and supplement the screened available keywords to the keyword list.
[0075] As a specific implementation of keyword processing, this step includes the following operations:
[0076] (1) If the keyword set has not been created, construct the keyword set;
[0077] (2) Sort the words already recorded in the second word segmentation dataset in descending order according to the TF-IDF values;
[0078] (3) Based on the sorting result, use the ordered integer set formed by the sorting serial numbers as the initial search set of the binary search algorithm, obtain the TF-IDF values of the words already recorded with serial numbers less than or equal to the specified truncation value, calculate the proportion of the sum of the obtained TF-IDF values of the words already recorded, use the proportion as the cumulative frequency index applied in the primary and secondary factor analysis method, combine the binary search algorithm and the primary and secondary factor analysis method to determine the available keywords in the second word segmentation dataset, and supplement the unrecorded available keywords to the keyword list;
[0079] Among them, the calculation formula for the proportion is:
[0080]
[0081] This step is based on the second word segmentation dataset, and combines the binary search method and the primary and secondary factor analysis method to determine the keyword vocabulary. If the keyword vocabulary has not been constructed yet, the initial construction is completed; according to the second word segmentation dataset, the recorded words in the data table are sorted in descending order according to the TF-IDF values; based on the descending order result of the TF-IDF values, the ordered integer set composed of the sorting serial numbers is used as the initial search set of the binary search algorithm, and the proportion of the sum of the TF-IDF of the recorded words with the serial number less than or equal to the specified truncation value is used as the cumulative frequency index applied in the primary and secondary factor analysis method. Combine the binary search algorithm and the primary and secondary factor analysis method to determine the available keywords in the second word segmentation dataset, and supplement the unrecorded available keywords to the keyword vocabulary. Specifically, the process of combining the binary search algorithm and the primary and secondary factor analysis method to determine the available keywords is as follows: set the cumulative frequency, and use the initial search set as the current search set; determine the middle serial number of the current search set as the current specified truncation value, and determine the proportion of the sum of the TF-IDF of the recorded words with the serial number less than or equal to the current specified truncation value according to the corresponding formula, and compare it with the set cumulative frequency; if the two are equal, use the recorded words in the sorting result with the sorting serial number less than or equal to the current specified truncation value as the available keywords and end the search; if the proportion is less than the set cumulative frequency, and the current search interval cannot be split in half, use the recorded words with the sorting serial number less than or equal to the right end serial number of the current search set as the available keywords and end the search, otherwise use the right half interval of the current search set as the current search set for the next round of search and continue the next round of search; if the proportion is greater than the set cumulative frequency, and the current search interval cannot be split in half, use the recorded words with the sorting serial number less than or equal to the current specified truncation value as the available keywords and end the search, otherwise use the left half interval of the current search set as the current search set for the next round of search and continue the next round of search;
[0082] Specifically, as Figure 3 shown, for the initial implementation process of the task, the keyword vocabulary can be constructed according to the method described in this step; for the subsequent cycle maintenance process of the task, according to the set maintenance cycle, determine the newly added available keywords based on the newly added recorded words in the re-constructed second word segmentation dataset according to the method described in this step, and complete the supplement and maintenance of the keyword vocabulary.
[0083] Step S400 Keyword Encoding: Encode the keywords in the keyword vocabulary through a word embedding model to obtain the corresponding encoding vectors, and supplement the keywords and the corresponding encoding vectors to the keyword set in the vector database.
[0084] In this step of this embodiment, based on the keyword vocabulary, a word embedding model is used to encode the keywords and store them in the keyword set in the vector database. If the keyword set is not constructed in the vector database, the initial construction is completed; the newly added available keywords in the keyword vocabulary are sequentially selected, the selected keywords are encoded using the word embedding model, and an appropriate index type is selected based on the task, and then the keywords and their corresponding encoded vectors and other data are stored in the constructed keyword set in the vector database.
[0085] Specifically, as Figure 3 shown, for the initial implementation process of the task, the keywords in the keyword vocabulary can be encoded as a whole and stored in the vector database according to the method described in this step; for the subsequent periodic maintenance process of the task, according to the set maintenance period, the keyword set in the vector database is supplemented and maintained based on the newly added available keywords in the keyword vocabulary according to the method described in this step.
[0086] Step S500 Keyword Recommendation: For a given word, the given word is encoded using a word embedding model to obtain the encoded vector corresponding to the given word. Based on the encoded vector corresponding to the given word, keywords with similarity matching are retrieved from the keyword set in the vector database through vector similarity search. For the retrieved words, according to the task requirements, the words that are exactly the same as the given word are deleted.
[0087] In this step of this embodiment, based on the keyword set in the vector database, keyword automatic recommendation is realized based on the given word through vector similarity retrieval. As Figure 3 shown, for a given word, it is encoded using a word embedding model to obtain the encoded vector corresponding to the given word, and it is used as the current retrieval vector. After setting the number of results to be returned for similarity search according to the task requirements, keywords with high semantic similarity to the given word are retrieved from the keyword set based on the vector database similarity search strategy, and the words that are exactly the same as the given word are removed according to the task requirements.
[0088] The method of this embodiment combines Chinese word segmentation, keyword generation, word embedding model technology, and vector database application technology, which can effectively extract keyword information in the corpus, deeply understand the semantics of keywords, and thus can efficiently and accurately implement the keyword automatic recommendation task.
[0089] Embodiment 2:
[0090] A semantic-based keyword automatic recommendation system of the present invention includes a data processing module, a word segmentation processing module, a keyword processing module, a keyword encoding module, and a keyword recommendation module.
[0091] The data processing module is used to perform the following: collect text data according to task requirements, store the text data in the first corpus, clean the text data to be cleaned in the first corpus, and add the cleaned text data to the second corpus.
[0092] As a specific implementation of the data processing module, the module is used to perform the following operations:
[0093] (1) determining the type and scope of text data to be collected according to business requirements, collecting text data based on the determined type and scope, and aggregating the collected text data into a first prediction database;
[0094] (2) regularly reading the text data to be cleaned in the first corpus through a program script, performing a data cleaning operation on the text data to be cleaned, and adding the cleaned text data to the second corpus, wherein the data cleaning operation includes removing duplicate data, removing irrelevant characters, and removing noise data.
[0095] During this step, if the first corpus and the second corpus have not been constructed, the initial construction is completed. According to the specific requirements of the task, the detailed types and collection scope of the text data to be collected are determined, and the text data collection is completed accordingly, and the data is aggregated into the first corpus. The data to be cleaned in the first corpus is read through the program script to perform a series of data cleaning processes, and the obtained text data is added to the second corpus. The data cleaning process of the text data includes removing duplicate text data, removing irrelevant characters, and removing noise data.
[0096] For the initial implementation process of the task, the first corpus and the second corpus can be directly constructed based on the initial collection of text data; for the subsequent periodic maintenance process of the task, the second corpus is supplemented and maintained according to the set maintenance cycle and based on the newly added data to be cleaned in the first corpus.
[0097] The word segmentation processing module is used to perform the following: perform word segmentation processing on the text data in the second corpus, update the statistical data corresponding to the recorded words in the first word segmentation data set according to the words obtained by the word segmentation processing, and add unrecorded words and the parts of speech and statistical data corresponding to the unrecorded words to the first word segmentation data set, screen the recorded words in the first word segmentation data set, and calculate the TF-IDFFTF-IDF values of the recorded words in combination with the scale of the text data in the second corpus, record the recorded words and the corresponding TF-IDF values in the second word segmentation data set, wherein, for the recorded words and the unrecorded words, the statistical data include the number of occurrences of the words and the number of texts containing the words.
[0098] As a specific implementation of the word segmentation processing module, this module is used to perform the following operations:
[0099] (1) Read the text data to be segmented in the second corpus in sequence. Call the Jieba segmentation algorithm through a program script, segment the read text data in the search mode of the segmentation mode, and remove stop words in combination with a preset stop word list during the segmentation process. There are two types of words obtained by segmentation, namely the words already recorded in the first segmented dataset and the words not recorded in the first segmented dataset. For the words already recorded in the first segmented dataset, update the statistical data corresponding to the words already recorded in the first segmented dataset. For the words not recorded in the segmented dataset, supplement the unrecorded words, the corresponding part-of-speech, and the statistical data to the first segmented dataset;
[0100] (2) For the first segmented dataset, overall screen all the words already recorded in the first segmented dataset according to the set maintenance period and based on the set target part-of-speech range, and read the statistical data corresponding to the screened words already recorded in sequence. Combine the scale information of the text data in the second corpus to calculate the TF-IDF of the read words already recorded, and record the words already recorded and their corresponding TF-IDF values in the second segmented dataset;
[0101] Among them, the calculation formula for the TF-IDF value of the words already recorded is as follows:
[0102]
[0103] In this step, segment the data to be segmented in the second corpus to determine the first segmented dataset, and obtain the second segmented dataset through target part-of-speech screening and by combining the improved TF-IDF method. If the first segmented dataset has not been constructed yet, complete the initial construction. Read the data to be segmented in the second corpus in sequence, use the Chinese word segmentation method to segment the read text data, and then based on the words obtained by segmentation, update the statistical data corresponding to the words already recorded in the first segmented dataset, and supplement the unrecorded words, the corresponding part-of-speech, and their corresponding statistical data to the first segmented dataset; for the first segmented dataset, overall screen all the words already recorded in this dataset according to the set target part-of-speech range, and read the statistical data corresponding to the screened words already recorded in sequence, and combine the scale information of the text data in the second corpus, calculate the TF-IDF of the read words already recorded according to the corresponding formula, and record the words already recorded and their corresponding TF-IDF values in the second segmented dataset.
[0104] Among them, the process of text data segmentation includes: calling the Jieba segmentation algorithm through a program script, segmenting the read text data in the search mode of the segmentation mode, and removing stop words in combination with a preset stop word list during the segmentation process to obtain the final segmented data.
[0105] Specifically, for the initial implementation process of the task, the first word segmentation dataset and the second word segmentation dataset can be determined in sequence; for the subsequent cycle maintenance process of the task, according to the set maintenance cycle, supplement and maintain the first word segmentation dataset based on the newly added data to be segmented in the second corpus, and reconstruct the second word segmentation dataset accordingly.
[0106] The keyword processing module is used to perform the following: screen available keywords from the second word segmentation dataset through the binary search method and the primary and secondary factor analysis method, and supplement the screened available keywords to the keyword vocabulary.
[0107] As a specific implementation of the keyword processing module, this module is used to perform the following operations:
[0108] (1) If the keyword set has not been created, construct the keyword set;
[0109] (2) Sort the recorded words in the second word segmentation dataset in descending order according to the TF-IDF values;
[0110] (3) Based on the sorting result, use the ordered integer set composed of the sorting serial numbers as the initial search set of the binary search algorithm, obtain the TF-IDF values of the recorded words with the serial numbers less than or equal to the specified truncation value, calculate the proportion of the sum of the obtained TF-IDF values of the recorded words, use the proportion as the cumulative frequency index applied in the primary and secondary factor analysis method, combine the binary search algorithm and the primary and secondary factor analysis method to determine the available keywords in the second word segmentation dataset, and supplement the unrecorded available keywords to the keyword vocabulary;
[0111] Among them, the calculation formula for the proportion is:
[0112]
[0113] This module determines the keyword vocabulary based on the second word segmentation dataset, combining the binary search method and the primary and secondary factor analysis method. If the keyword vocabulary has not been constructed yet, it completes the initial construction. According to the second word segmentation dataset, the words recorded in the data table are sorted in descending order according to their TF-IDF values. Based on the descending order result of the TF-IDF values, the ordered integer set formed by the sorting serial numbers is used as the initial search set for the binary search algorithm, and the proportion of the sum of the TF-IDF values of the recorded words with serial numbers less than or equal to the specified truncation value is used as the cumulative frequency index applied in the primary and secondary factor analysis method. Combining the binary search algorithm and the primary and secondary factor analysis method, the available keywords in the second word segmentation dataset are determined, and the unrecorded available keywords are added to the keyword vocabulary. Specifically, the process of determining the available keywords by combining the binary search algorithm and the primary and secondary factor analysis method is as follows: Set the cumulative frequency, and use the initial search set as the current search set. Determine the middle serial number of the current search set as the current specified truncation value, calculate the proportion of the sum of the TF-IDF values of the recorded words with serial numbers less than or equal to the current specified truncation value according to the corresponding formula, and compare it with the set cumulative frequency. If the two are equal, the recorded words in the sorting result with serial numbers less than or equal to the current specified truncation value are used as the available keywords, and the search ends. If the proportion is less than the set cumulative frequency, and the current search interval cannot be split in half anymore, the recorded words with serial numbers less than or equal to the right end serial number of the current search set are used as the available keywords, and the search ends. Otherwise, the right half interval of the current search set is used as the current search set for the next round of search, and the next round of search continues. If the proportion is greater than the set cumulative frequency, and the current search interval cannot be split in half anymore, the recorded words with serial numbers less than or equal to the current specified truncation value are used as the available keywords, and the search ends. Otherwise, the left half interval of the current search set is used as the current search set for the next round of search, and the next round of search continues.
[0114] Specifically, for the initial implementation process of the task, a keyword vocabulary can be constructed. For the subsequent periodic maintenance process of the task, based on the set maintenance period and taking the newly added recorded words in the newly constructed second word segmentation dataset as the basis, the newly added available keywords are determined, and the keyword vocabulary is supplemented and maintained.
[0115] The keyword encoding module is used to perform the following: Encode the keywords in the keyword vocabulary through a word embedding model to obtain the corresponding encoding vectors, and add the keywords and the corresponding encoding vectors to the keyword set in the vector database.
[0116] In this embodiment, this module is used to encode keywords based on a keyword vocabulary using a word embedding model and store them in the keyword set in the vector database. If the keyword set is not constructed in the vector database, the initial construction is completed; the newly added available keywords in the keyword vocabulary are sequentially selected, the selected keywords are encoded using the word embedding model, and an appropriate index type is selected based on the task, and then data such as the keywords and their corresponding encoded vectors are stored in the constructed keyword set in the vector database.
[0117] Specifically, for the initial implementation process of the task, the keywords in the keyword vocabulary can be encoded as a whole and stored in the vector database; for the subsequent periodic maintenance process of the task, based on the set maintenance period and taking the newly added available keywords in the keyword vocabulary as the basis, the keyword set in the vector database is supplemented and maintained.
[0118] The keyword recommendation module is used to perform the following: for a given word, the given word is encoded using a word embedding model to obtain the encoded vector corresponding to the given word, and based on the encoded vector corresponding to the given word, keywords with similarity matching are retrieved from the keyword set in the vector database by means of vector similarity search. For the retrieved words, the words that are exactly the same as the given word are deleted according to the task requirements.
[0119] This module in this embodiment is used to automatically recommend keywords based on a given word through vector similarity retrieval based on the keyword set in the vector database. For a given word, it is encoded using a word embedding model to obtain the encoded vector corresponding to the given word, and this is used as the current retrieval vector. After setting the number of results to be returned for the similarity search according to the task requirements, keywords with high semantic similarity to the given word are retrieved from the keyword set based on the vector database similarity search strategy, and the words that are exactly the same as the given word are removed according to the task requirements.
[0120] This system can execute the method disclosed in Embodiment 1 to achieve automatic keyword recommendation.
[0121] The above has introduced in detail the method and system for automatic keyword recommendation based on semantics provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A semantic-based automatic keyword recommendation method, characterized in that, The steps include: Data processing: collect text data according to task requirements, store the text data in the first corpus, clean the text data to be cleaned in the first corpus, and add the cleaned text data to the second corpus; Word segmentation processing: perform word segmentation processing on the text data in the second corpus, update the statistical data corresponding to the recorded words in the first word segmentation data set according to the words obtained by the word segmentation processing, and add unrecorded words and the parts of speech and statistical data corresponding to the unrecorded words to the first word segmentation data set, screen the recorded words in the first word segmentation data set, and calculate the TF-IDFFTF-IDF values of the recorded words in combination with the scale of the text data in the second corpus, and record the recorded words and the corresponding TF-IDF values in the second word segmentation data set, where, for the recorded words and unrecorded words, the statistical data include the number of occurrences of the words and the number of texts containing the words; Keyword processing: Use binary search and primary and secondary factor analysis to select available keywords from the second segmentation data set, and add the selected available keywords to the keyword word list; Keyword encoding: Encode the keywords in the keyword vocabulary through the word embedding model to obtain the corresponding encoding vector, and add the keywords and the corresponding encoding vector to the keyword set of the vector database; Keyword recommendation: For a given word, encode the given word through the word embedding model to obtain the encoding vector corresponding to the given word. Based on the encoding vector corresponding to the given word, retrieve keywords with similarity matching from the keyword set of the vector database through vector similarity search. For the retrieved words, delete the words that are exactly the same as the given words according to task requirements.
2. The semantic-based keyword automatic recommendation method according to claim 1, wherein Data processing includes the following steps: According to business needs, determine the type and scope of text data to be collected, collect text data based on the determined type and scope, and aggregate the collected text data into a first forecast library; The text data to be cleaned in the first corpus is read regularly by a program script, a data cleaning operation is performed on the text data to be cleaned, and the cleaned text data is added to the second corpus, wherein the data cleaning operation includes removing duplicate data, removing irrelevant characters, and removing noise data.
3. The semantic-based keyword automatic recommendation method according to claim 1, wherein Word segmentation includes the following steps: The text data to be segmented in the second corpus is read in sequence, the Jieba segmentation algorithm is called through a program script, the read text data is segmented when the segmentation mode is the search mode, and stop words are removed in combination with a preset stop word list during the segmentation process, and there are two types of words obtained by the segmentation, namely, words recorded in the first segmentation data set and words not recorded in the first segmentation data set, for the words recorded in the first segmentation data set, the statistical data corresponding to the recorded words in the first segmentation data set are updated, and for the words not recorded in the segmentation data set, the unrecorded words and the part of speech and statistical data corresponding to the unrecorded words are added to the first segmentation data set; For the first word segmentation dataset, according to the set maintenance period, all the recorded words in the first word segmentation dataset are screened as a whole based on the set target part-of-speech range, and the statistical data corresponding to the recorded words obtained by the screening are read in sequence. Combining with the scale information of the text data in the second corpus, calculate the TF-IDF of the recorded words read, and record the recorded words and their corresponding TF-IDF values in the second word segmentation dataset; Among them, the calculation formula for the TF-IDF value of the recorded word is as follows:
4. The semantic-based keyword automatic recommendation method according to claim 1, wherein The keyword processing includes the following steps: If the keyword set has not been created, construct the keyword set; Sort the recorded words in the second word segmentation dataset in descending order according to the TF-IDF value; Based on the sorting result, use the ordered integer set composed of the sorting serial numbers as the initial search set of the binary search algorithm, obtain the TF-IDF values of the recorded words whose serial numbers are less than or equal to the specified truncation value, calculate the ratio of the sum of the obtained TF-IDF values of the recorded words, and use the ratio as the cumulative frequency index applied in the primary and secondary factor analysis method. Combine the binary search algorithm and the primary and secondary factor analysis method to determine the available keywords in the second word segmentation dataset, and supplement the unrecorded available keywords to the keyword list; Among them, the calculation formula for the ratio is:
5. A semantic-based keyword automatic recommendation system, characterized in that, Used to implement keyword automatic recommendation through a semantic-based keyword automatic recommendation method as described in any one of claims 1-4. The system includes a data processing module, a word segmentation processing module, a keyword processing module, a keyword encoding module, and a keyword recommendation module; The data processing module is used to perform the following: collect text data according to the task requirements, store the text data in the first corpus, and perform data cleaning on the text data to be cleaned in the first corpus, and supplement the cleaned text data to the second corpus; The word segmentation processing module is used to perform the following: perform word segmentation processing on the text data in the second corpus, update the statistical data corresponding to the recorded words in the first word segmentation dataset according to the words obtained by the word segmentation processing, and supplement the unrecorded words and the corresponding part-of-speech and statistical data to the first word segmentation dataset. Screen the recorded words in the first word segmentation dataset, and combine with the scale of the text data in the second corpus to calculate the TF-IDF value of the recorded words. Record the recorded words and the corresponding TF-IDF values in the second word segmentation dataset. Among them, for the recorded words and unrecorded words, the statistical data includes the number of occurrences of the word and the number of texts containing the word; The keyword processing module is used to perform the following: screen available keywords from the second word segmentation dataset through the binary search method and the primary and secondary factor analysis method, and supplement the screened available keywords to the keyword list; The keyword encoding module is used to perform the following: encode the keywords in the keyword list through the word embedding model to obtain the corresponding encoding vectors, and supplement the keywords and the corresponding encoding vectors to the keyword set in the vector database; The keyword recommendation module is used to perform the following: For a given word, encode the given word through a word embedding model to obtain an encoded vector corresponding to the given word, retrieve keywords with similarity matching from the keyword set in the vector database based on the encoded vector corresponding to the given word by means of vector similarity search, and for the retrieved words, delete the words that are exactly the same as the given word according to the task requirements.
6. The semantic-based keyword automatic recommendation system according to claim 5, characterized in that, The data processing module is used to perform the following operations: According to the business requirements, determine the type and scope of the text data to be collected, collect the text data based on the determined type and scope, and gather the collected text data into the first corpus; Regularly read the text data to be cleaned in the first corpus through a program script, perform data cleaning operations on the text data to be cleaned, and append the cleaned text data to the second corpus, where the data cleaning operations include removing duplicate data, removing irrelevant characters, and removing noise data.
7. The semantic-based keyword automatic recommendation system according to claim 5, characterized in that The word segmentation processing module is used to perform the following operations: Read the text data to be segmented in the second corpus in sequence, call the Jieba word segmentation algorithm through a program script, segment the read text data in the case of the search segmentation mode, and remove stop words in combination with a preset stop word list during the word segmentation process. There are two types of words obtained by word segmentation, namely words already recorded in the first segmented dataset and words not recorded in the first segmented dataset. For words already recorded in the first segmented dataset, update the statistical data corresponding to the words already recorded in the first segmented dataset. For words not recorded in the segmented dataset, append the unrecorded words and their corresponding part-of-speech and statistical data to the first segmented dataset; For the first segmented dataset, based on the set maintenance period and the set target part-of-speech range, overall screen all the words already recorded in the first segmented dataset, read the statistical data corresponding to the screened recorded words in sequence, and calculate the TF-IDF of the read recorded words in combination with the scale information of the text data in the second corpus. Record the recorded words and their corresponding TF-IDF values in the second segmented dataset; Among them, the calculation formula for the TF-IDF value of a recorded word is as follows:
8. The semantic-based keyword automatic recommendation system according to claim 5, characterized in that The keyword processing module is used to perform the following: If the keyword set has not been created, construct the keyword set; Sort the recorded words in the second segmented dataset in descending order according to the TF-IDF value; Based on the sorting result, use the ordered integer set composed of the sorting serial numbers as the initial search set of the binary search algorithm, obtain the TF-IDF values of the recorded words with the serial number less than or equal to the specified truncation value, calculate the proportion of the sum of the obtained TF-IDF values of the recorded words, use the proportion as the cumulative frequency index applied in the primary and secondary factor analysis method, and determine the available keywords in the second segmented dataset in combination with the binary search algorithm and the primary and secondary factor analysis method, and append the unrecorded available keywords to the keyword list; Among them, the calculation formula for the proportion is: