Search word association recommendation method and system, electronic equipment and storage medium
By constructing a statistical causal probability map and Bayesian network, the problem of failure to fully utilize the statistical causal relationship between search terms in the existing technology is solved, and a more accurate and reasonable search term recommendation effect is achieved.
Patent Information
- Application Number
- CN202510032615.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art fails to make full use of the statistical causal relationship between search terms in the recommendation of search terms, resulting in poor recommendation results.
Construct the statistical causal probability graph of all text vocabulary in a given field, establish directed edges between words based on the Bayesian network construction program, generate the search word association probability graph, and realize the associative recommendation of search word.
By making full use of the statistical characteristics of the text itself and the word co-occurrence relationship, the accuracy and rationality of search term recommendations are improved, and it does not depend on the user's search habits.
Smart Images

Figure CN120030214A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a search term association recommendation method, system, electronic device and storage medium. Background Art
[0002] With the development of Internet technology, in order to meet the needs of Internet users for obtaining information in different fields, search engines receive search keywords entered by users, usually associate the search keywords, and recommend related associated words to users. Given all the texts in a field, it is an important requirement to provide users with reasonable search term association recommendations when they search. In most of the current search engines, the default search category is web search, which refers to searching all network resources related to the terms that the user intends to search for based on the user's search request. In order to avoid users entering long search terms, search engines provide users with a search prompt function, that is, when the user enters the prefix or suffix of the search term, the search engine prompts the user with multiple commonly used search terms, and the user can directly perform an online search for commonly used search terms by selecting the confirmation operation.
[0003] Search engine (English name: search engine: according to the usage scenario, user needs and certain algorithms, a specific strategy is used to retrieve specified information from the data and feed it back to the user, providing users with fast and highly relevant information services. Natural language processing: an important research direction in the field of artificial intelligence, integrating knowledge from multiple disciplines such as linguistics, computer science, machine learning, mathematics, cognitive psychology, etc. It is an interdisciplinary subject integrating computer science, artificial intelligence and linguistics. It includes two main aspects: natural language understanding and natural language generation. The research content includes multiple levels such as characters, words, phrases, sentences, paragraphs and chapters. It is a bridge for communication between machine language and human language. It aims to enable machines to understand, interpret and generate human language, realize effective communication between humans and machines, and enable computers to perform tasks such as language translation, sentiment analysis, and text summarization.
[0004] Common solutions in the prior art are as follows: 1) Counting high-frequency words in historical searches, based on the search information input by the user, for example, the user's search words (if it is a search statement, the search words are first segmented to obtain the search words), calculate the similarity between the search words and the high-frequency words in historical searches for analysis; 2) First search for corresponding text fragments based on the search words, and find associative recommended words for the current search words in these text fragments; 3) Performing associative recommendations for search words (for example, the user's high-frequency search words) based on the user's search behavior habits.
[0005] Among the above methods, methods 1) and 2) mainly utilize the similarity and co-occurrence features between search terms, and method 3 only analyzes search behavior habits. These methods do not start from the perspective of statistical language models and fail to fully utilize the statistical causal relationship (conditional probability graph) between search terms to make reasonable recommendations.
[0006] Therefore, the applicant proposes a solution to construct a statistical causal probability graph of all text words in a given field, and then implements an associative recommendation method for search terms based on the statistical causal probability graph of the words to solve the problems existing in the prior art. Summary of the invention
[0007] The present invention proposes a search term associative recommendation method, system, electronic device and storage medium, which solves the problem that the search term recommendation method in the prior art fails to make full use of the statistical causal relationship between search terms to make reasonable recommendations. The technical solution of the present invention is implemented as follows:
[0008] A search word association recommendation method includes the following steps: constructing a search word association probability graph, and forming a search word generation scheme based on the generated search word association probability graph;
[0009] The search term association probability graph is constructed by the following steps: all given texts are divided into blocks, and the contents of all text blocks A1, A2, A3…, An are summarized into texts B1, B2, B3,…, Bn within 50 words; all text blocks A1, A2, A3,…, An are segmented to obtain segmentation sets P1, P2, P3,…, Pn respectively, let P = P1∪P2…∪Pn, then P is the dictionary of the content of these given texts, and the first K items are taken according to the word frequency, recorded as P(K); all text summaries B1, B2, B3,…, Bn are segmented to obtain segmentation sets Q1, Q2,…, Qn respectively, let Q = Q1∪Q2…∪Qn, then Q is the dictionary of the text summaries of these given texts, and the first L items are taken according to the word frequency, recorded as Q(L); let R = P(K)∪Q(L), and take R as the search term candidate dictionary of the given text;
[0010] Assume that all the words in the search term candidate dictionary R are W1, W2, ..., Wm, a total of m, construct a word co-occurrence statistics table, record the co-occurrence records of W1, W2, ..., Wm in Pi∪Qi (i = 1, 2, 3, ... n), there are n records in total; form a word co-occurrence table T(n, m) with n rows and m columns; the word Wj in each column of the word co-occurrence table T(n, m) represents a node (j = 1, 2, ..., m); run the Bayesian network construction program, establish the Bayesian probability network BN between Wj, and use the directed edges of the constructed Bayesian probability network BN to obtain the adjacency relationship matrix BNW of each word W1, W2, ..., Wm in the search term candidate dictionary R; input the search statement or search term, and generate a search term association scheme based on the search term association probability graph constructed in the above steps.
[0011] As a further technical solution, the search term generation solution specifically includes the following steps:
[0012] 1) Enter the search statement or search term;
[0013] 2) Segment the user's input I to obtain a word set IW;
[0014] 3) Define the word set MIS = IW∩R, and assume that there are s words in MIS, which are I1, I2, I3, ..., Is;
[0015] 4) Based on the vocabulary Ik (k = 1, 2, ..., s) in MIS, search for the adjacent word set E(Ik) of Ik in the adjacency matrix BNW, (k = 1, 2, ..., s);
[0016] 5) Find the union NIS = E(I1)∪E(I2)∪…∪E(Is), where NIS is the set of candidate search association words for input I. Suppose there are t words in NIS, namely O1, O2,…, Ot;
[0017] 6) Perform reasoning based on the Bayesian probability network BN and calculate respectively:
[0018] P(O1|I1=1, I2=1, ..., Is=1, other words in the search word candidate dictionary R=0),
[0019] P(O2|I1=1,I2=1,…,Is=1,other words in the search word candidate dictionary R=0),
[0020] P(O3|I1=1,I2=1,…,Is=1,other words in search term candidate dictionary R=0),…,
[0021] P(Ot|I1=1, I2=1, ..., Is=1, other words in the search term candidate dictionary R=0);
[0022] 7) Take the first J largest probabilities among the above t probabilities, and the corresponding words are the search association words of input I.
[0023] As a further technical solution, the Bayesian probability network BN is constructed by calculating the scoring function method. The method is as follows: by calculating the maximum likelihood estimate and penalty term of the model, that is, the product of the number of model parameters and the logarithm of the number of samples, the goodness of fit and complexity of the model are comprehensively measured; its mathematical expression is
[0024] BIC=-2*log(L)+k*log(n)
[0025] Where L is the maximum likelihood of the model, k is the number of model parameters, and n is the number of samples.
[0026] As a further technical solution, the Bayesian probability network BN is constructed using the hill climbing method as follows:
[0027] 1) Initialization: Start from an initial solution or a random solution in the solution space.
[0028] 2) Generate neighborhood solutions: Based on the current solution, generate several neighborhood solutions. The neighborhood solutions are solutions obtained by making small adjustments to the current solution.
[0029] 3) Evaluate and select: Evaluate the quality of each neighborhood solution, usually using a scoring function to measure the quality of the solution, and select the neighborhood solution with the best score as the new current solution.
[0030] 4) Update: Take the selected optimal neighborhood solution as the current solution and continue to generate new neighborhood solutions starting from this new solution.
[0031] 5) Stop condition: If there is no better neighborhood solution (that is, the scores of all neighborhood solutions are lower than the current solution), the search is stopped and the current solution is the local optimal solution.
[0032] As a further technical solution, text segmentation is implemented using an extractive summary model, a generative summary model or a large model.
[0033] A search word association recommendation system, the search word association recommendation system comprises: a receiving module, an acquisition module and a recommendation module; wherein the receiving module is used to receive the search word input by the user;
[0034] The acquisition module is used to acquire an associative word corresponding to the search word in an associative word library according to a search word associative recommendation method as described in any one of claims 1 to 7 above;
[0035] The recommendation module is used to recommend the associated word corresponding to the search word to the user.
[0036] An electronic device for search term association recommendation, comprising one or more processors;
[0037] a memory for storing one or more programs,
[0038] When the one or more programs are executed by the one or more processors, the one or more processors implement the search term association recommendation method as described in any one of claims 1 to 5.
[0039] A non-temporary storage medium, when the instructions in the storage medium are executed by the above-mentioned search word association recommendation electronic device, enables the device to execute the above-mentioned search word association recommendation method.
[0040] Compared with the prior art, this solution has the following beneficial effects:
[0041] 1) It does not rely on the user's search habits and fully considers the statistical characteristics of the text itself.
[0042] 2) By using text chunking, the context window of word co-occurrence can be determined more reasonably.
[0043] 3) On the basis of text segmentation, the text segmentation formed after summarization is added. These segmentations are usually closer to the input search terms, which improves the rationality of subsequent search term reasoning. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0045] Figure 1 The present invention is a method flow chart of a search term association recommendation method. DETAILED DESCRIPTION
[0046] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0047] Given all the text in a field, it is an important requirement to provide reasonable search term association recommendations for users when they search. The conventional ideas in the prior art are as follows:
[0048] 1. Count the high-frequency words in historical searches. Based on the search information entered by the user, for example, the user's search words (if it is a search statement, the search words are first segmented to obtain the search words), calculate the similarity between the search words and the high-frequency words in historical searches for analysis;
[0049] 2. First, search for corresponding text fragments based on the search term, and find the associated recommended words of the current search term in these text fragments;
[0050] 3. Make associative recommendations for search terms (e.g., the user's high-frequency search terms) based on the user's search behavior habits.
[0051] Among the above methods, methods 1 and 2 mainly utilize the similarity and co-occurrence features between search terms, and method 3 only analyzes search behavior habits. These methods do not start from the perspective of statistical language models and fail to fully utilize the statistical causal relationship (conditional probability graph) between search terms to make reasonable recommendations.
[0052] In view of the above problems, the applicant proposed a search term association recommendation method, referring to Figure 1 ,The specific implementation of the present invention is mainly divided into two parts: 1. ,Scheme for constructing the search word association probability graph; 2. ,Scheme for generating the search word association.
[0053] (I) Construction scheme of search term association probability map
[0054] 1. Divide all given text into blocks. Commonly used block division schemes are:
[0055] (1) Fixed character length text segmentation: divide the text into several blocks according to the preset fixed character length. (2) Fixed word sentence number text segmentation: divide the text into several blocks according to the preset fixed sentence number. (3) Text segmentation method based on BERT In order to enable the BERT model to learn the relationship between two sentences, a binary classification task is designed during its pre-training process: input two sentences into BERT at the same time and predict whether the second sentence is the next sentence of the first sentence. When segmenting, a sliding window method is used to input two adjacent sentences into the BERT model for binary classification prediction. If the prediction score is low, it means that the semantic relationship between the two sentences is weak, and this can be used as the segmentation point of the text. (4) Segmentation based on the title structure: the text between the titles is taken as a block. After this step, all text blocks A1, A2, A3, …, An are obtained, with a total of n blocks.
[0056] 2. Summarize the contents of all text blocks A1, A2, A3, ..., An into texts B1, B2, B3, ..., Bn of less than 50 words. This step can be implemented using an extractive summarization model, a generative summarization model, or a large model.
[0057] 3. Perform word segmentation on all text blocks A1, A2, A3, ..., An, and obtain word segmentation sets P1, P2, P3, ..., Pn respectively. Note: The word segmentation described in this article has taken into account the use of customized dictionaries and stop dictionaries, which will not be repeated here.
[0058] 4. Segment all the text summaries B1, B2, B3, ..., Bn to obtain the segmentation sets Q1,
[0059] Q2,…,Qn.
[0060] 5. Let P = P1∪P2…∪Pn, then P is the dictionary of the text content of these given texts. Take the first K items according to the frequency of the words, and record them as P(K).
[0061] 6. Let Q = Q1∪Q2…∪Qn, then Q is the dictionary summarizing the given texts. Take the first L items according to the frequency of the words, and record them as Q(L).
[0062] 7. Let R = P(K) ∪ Q(L) and take R as the candidate dictionary of search terms for the given text.
[0063] 8. Let all the words in the search term candidate dictionary R be W1, W2, ..., Wm, a total of m. Construct a word co-occurrence statistics table to record the records of W1, W2, ..., Wm co-occurring (appearing together) in Pi∪Qi (i=1,2,3, ...n), a total of n records. Form a word co-occurrence table T(n,m) with n rows and m columns.
[0064] 9. In the word co-occurrence table T(n,m), each word Wj in each column represents a node (j=1,2,…,m). Run the Bayesian network construction program to build a Bayesian probability network BN between Wj. The specific construction method is as follows:
[0065] (1) Calculate the scoring function BIC (Bayesian information criterion). BIC is a comprehensive measure of the goodness of fit and complexity of the model by calculating the maximum likelihood estimate of the model and the penalty term (the product of the number of model parameters and the logarithm of the number of samples). Its mathematical expression is BIC = -2
[0066] *log(L)+k*log(n), where L is the maximum likelihood of the model, k is the number of model parameters, and n is the number of samples.
[0067] (2) Use the Hill-Climbing method (HC) to get the minimum value of BIC. The basic process of HC is as follows:
[0068] 1) Initialization: Start from an initial solution (or a random solution in the solution space).
[0069] 2) Generate neighborhood solutions: Generate several neighborhood solutions based on the current solution. The neighborhood solution is obtained by making small adjustments to the current solution.
[0070] 3) Evaluate and select: Evaluate the quality of each neighborhood solution, usually using a scoring function to measure the quality of the solution. Select the neighborhood solution with the best score as the new current solution.
[0071] 4) Update: Take the selected optimal neighborhood solution as the current solution and continue to generate new neighborhood solutions starting from this new solution.
[0072] 5) Stop condition: If there is no better neighborhood solution (that is, the scores of all neighborhood solutions are lower than the current solution), the search is stopped and the current solution is the local optimal solution.
[0073] 10. Using the directed edges of the constructed Bayesian probability network BN, the adjacency relationship matrix BNW of each word W1, W2, ..., Wm in the search term candidate dictionary R is obtained.
[0074] (II) Search term association generation scheme
[0075] 1. Enter a search statement or search term.
[0076] 2. Segment the user’s input I to obtain a word set IW.
[0077] 3. Define the word set MIS = IW∩R. Suppose there are s words in MIS, namely I1, I2, I3, …,
[0078] Is.
[0079] 4. Based on the vocabulary Ik (k = 1, 2, ..., s) in MIS, search for the adjacent word set E (Ik) of Ik in the adjacency matrix BNW (k = 1, 2, ..., s).
[0080] 5. Find the union NIS = E(I1)∪E(I2)∪…∪E(Is). NIS is the set of candidate search association words for input I. Suppose there are t words in NIS, namely O1, O2,…, Ot.
[0081] 6. According to the Bayesian probability network BN, we can calculate
[0082] P(O1|I1=1,I2=1,…,Is=1, other words in the search term candidate dictionary R=0),
[0083] P(O2|I1=1,I2=1,…,Is=1, other words in the search term candidate dictionary R=0),
[0084] P(O3|I1=1,I2=1,…,Is=1,other words in the search term candidate dictionary R=0),…,
[0085] P(Ot|I1=1,I2=1,…,Is=1, other words in the search term candidate dictionary R=0).
[0086] 7. Take the first J largest probabilities among the above t probabilities, and the corresponding words are the search association words of input I.
[0087] Compared with the prior art, this solution has the following beneficial effects:
[0088] 1) It does not rely on the user's search habits and fully considers the statistical characteristics of the text itself.
[0089] 2) By using text chunking, the context window of word co-occurrence can be determined more reasonably.
[0090] 3) On the basis of text segmentation, the text segmentation formed after summarization is added. These segmentations are usually closer to the input search terms, which improves the rationality of subsequent search term reasoning.
[0091] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A search term association recommendation method, characterized in that: The following steps are involved: Constructing a search term association probability map, and forming a search term generation scheme based on the generated search term association probability map; The search term association probability graph is constructed by the following steps: all given texts are divided into blocks, and the contents of all text blocks A1, A2, A3…, An are summarized into texts B1, B2, B3,…, Bn within 50 words; all text blocks A1, A2, A3,…, An are segmented to obtain segmentation sets P1, P2, P3,…, Pn respectively, let P = P1∪P2…∪Pn, then P is the dictionary of the content of these given texts, and the first K items are taken according to the word frequency, recorded as P(K); all text summaries B1, B2, B3,…, Bn are segmented to obtain segmentation sets Q1, Q2,…, Qn respectively, let Q = Q1∪Q2…∪Qn, then Q is the dictionary of the text summaries of these given texts, and the first L items are taken according to the word frequency, recorded as Q(L); let R = P(K)∪Q(L), and take R as the search term candidate dictionary of the given text; Assume that all the words in the search term candidate dictionary R are W1, W2, ..., Wm, a total of m, construct a word co-occurrence statistics table, record the co-occurrence records of W1, W2, ..., Wm in Pi∪Qi (i = 1, 2, 3, ... n), there are n records in total; form a word co-occurrence table T(n, m) with n rows and m columns; the word Wj in each column of the word co-occurrence table T(n, m) represents a node (j = 1, 2, ..., m); run the Bayesian network construction program, establish the Bayesian probability network BN between Wj, and use the directed edges of the constructed Bayesian probability network BN to obtain the adjacency relationship matrix BNW of each word W1, W2, ..., Wm in the search term candidate dictionary R; input the search statement or search term, and generate a search term association scheme based on the search term association probability graph constructed in the above steps.
2. A search term association recommendation method as claimed in claim 1, characterized in that: The search term generation scheme specifically includes the following steps: 1) Enter the search statement or search term; 2) Segment the user's input I to obtain a word set IW; 3) Define the word set MIS = IW∩R, and assume that there are s words in MIS, which are I1, I2, I3, ..., Is; 4) Based on the vocabulary Ik (k = 1, 2, ..., s) in MIS, search for the adjacent word set E(Ik) of Ik in the adjacency matrix BNW, (k = 1, 2, ..., s); 5) Find the union NIS = E(I1)∪E(I2)∪…∪E(Is), where NIS is the set of candidate search association words for input I. Suppose there are t words in NIS, namely O1, O2,…, Ot; 6) Perform reasoning based on the Bayesian probability network BN and calculate respectively: P(O1|I1=1, I2=1, ..., Is=1, other words in the search word candidate dictionary R=0), P(O2|I1=1,I2=1,…,Is=1,other words in the search word candidate dictionary R=0), P(O3|I1=1,I2=1,…,Is=1,other words in search term candidate dictionary R=0),…, P(Ot|I1=1, I2=1, ..., Is=1, other words in the search term candidate dictionary R=0); 7) Take the first J largest probabilities among the above t probabilities, and the corresponding words are the search association words of input I.
3. A search term association recommendation method as claimed in claim 1, characterized in that: The Bayesian probability network BN is specifically constructed by calculating the scoring function method, which is as follows: by calculating the maximum likelihood estimate and penalty term of the model, that is, the product of the number of model parameters and the logarithm of the number of samples, the goodness of fit and complexity of the model are comprehensively measured; its mathematical expression is BIC=-2*log(L)+k*log(n) Where L is the maximum likelihood of the model, k is the number of model parameters, and n is the number of samples.
4. A search term association recommendation method as claimed in claim 1, characterized in that: The Bayesian probability network BN is specifically constructed using the hill climbing method, and the method is as follows: 1) Initialization: Start from an initial solution or a random solution in the solution space. 2) Generate neighborhood solutions: Based on the current solution, generate several neighborhood solutions. The neighborhood solutions are solutions obtained by making small adjustments to the current solution. 3) Evaluate and select: Evaluate the quality of each neighborhood solution, usually using a scoring function to measure the quality of the solution, and select the neighborhood solution with the best score as the new current solution. 4) Update: Take the selected optimal neighborhood solution as the current solution and continue to generate new neighborhood solutions starting from this new solution. 5) Stop condition: If there is no better neighborhood solution (that is, the scores of all neighborhood solutions are lower than the current solution), the search is stopped and the current solution is the local optimal solution.
5. The search term association recommendation method according to claim 1, characterized in that: The text segmentation is implemented using an extractive summary model, a generative summary model or a large model.
6. A search term association recommendation system, characterized in that: The search term association recommendation system comprises: a receiving module, an acquisition module and a recommendation module; wherein the receiving module is used to receive the search term input by the user; The acquisition module is used to acquire an associative word corresponding to the search word in an associative word library according to a search word associative recommendation method as described in any one of claims 1 to 7 above; The recommendation module is used to recommend the associated word corresponding to the search word to the user.
7. An electronic device for search word association recommendation, characterized in that: comprising one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the search term association recommendation method as described in any one of claims 1 to 5.
8. A non-temporary storage medium, characterized in that: When the instructions in the storage medium are executed by a search term association recommendation electronic device as described in claim 7, the device is enabled to execute any one of the search term association recommendation methods in claims 1 to 5 above.