Search processing method, device, equipment and storage medium based on dictionary tree

By using a dictionary tree-based search processing method in search engines, the problem of low search request processing efficiency in the prior art is solved, and fast and effective matching and sorting of search suggestions words is achieved, and the overall search processing efficiency is improved.

CN111460311BActive Publication Date: 2025-05-16TENCENT CLOUD COMPUTING (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010479380.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-12
Filing Date
2020-05-29
Publication Date
2025-05-16
Estimated Expiration
2040-05-29

AI Technical Summary

Technical Problem

When processing search requests, existing search engines need to analyze a large amount of data, resulting in large amounts of computing resources and low reading efficiency of search data.

Method used

The search processing method based on the dictionary tree is adopted to find the corresponding character nodes in the pre-constructed dictionary tree by obtaining search characters, traversing the child node sequence, extracting the string and weights, and sorting and displaying the search suggested words according to the weights.

Benefits of technology

It improves the reading efficiency and search processing efficiency of search data, reduces the need to calculate the weight of search suggestions, and reduces the use of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111460311B_ABST
    Figure CN111460311B_ABST
Patent Text Reader

Abstract

This application relates to a search processing method, apparatus, device, and storage medium based on a trie. The method includes: acquiring a search character; searching for a corresponding character node in a pre-constructed trie based on the character sequence of the search character; traversing the child node sequence corresponding to the character node; extracting the string corresponding to the child node sequence from the trie; obtaining the weight of each string from the corresponding node of the string; generating search suggestion terms based on the search character and each string; determining the weight of each string as the weight of the corresponding search suggestion term; sorting the search suggestion terms according to the weights; and displaying the sorted search suggestion terms through a search page on a terminal. This method can effectively improve the storage and retrieval efficiency of business data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the domestic priority of the Chinese patent application filed with the China Patent Office on November 12, 2019, with application number: 2019111023461, and named “Search processing method, device, equipment and storage medium based on dictionary tree”, all contents of which are incorporated by reference in this application. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to a search processing method, device, equipment and storage medium based on a dictionary tree. Background Art

[0003] With the rapid development of Internet technology, the demand for searching input content using input boxes is growing. Some search engines and input methods have emerged, which can use dictionary trees to perform fuzzy matching based on the string entered by the user and push the matching results to the user. However, the current search result recommendation method still needs to analyze a large amount of data during the search process, and arrange the search results according to the calculated probability such as search frequency. The computing resources are large in the process of search word recommendation, and the reading efficiency of search data is low. Summary of the invention

[0004] It is necessary to provide a search processing method, device, equipment and storage medium based on a dictionary tree to address the technical problems of low data storage efficiency and low reading efficiency.

[0005] A search processing method based on a dictionary tree, the method comprising:

[0006] Get the search character;

[0007] According to the character sequence of the search character, a corresponding character node is searched in a pre-constructed dictionary tree, a sub-node sequence corresponding to the character node is traversed, a character string corresponding to the sub-node sequence is extracted from the dictionary tree, and a weight corresponding to each character string is obtained from a node corresponding to the character string;

[0008] Generating search suggestion words according to the search characters and each character string, and determining the weight corresponding to each character string as the weight of the corresponding search suggestion word;

[0009] The search suggestion words are sorted according to the weights, and the sorted search suggestion words are displayed on a search page of the terminal.

[0010] A search processing device based on a dictionary tree, the device comprising:

[0011] A data acquisition module, used to acquire search characters;

[0012] A data search module, used to search for corresponding character nodes in a pre-constructed dictionary tree according to the character sequence of the search character, traverse the sub-node sequence corresponding to the character node, extract the character string corresponding to the sub-node sequence from the dictionary tree, and obtain the weight corresponding to each character string from the node corresponding to the character string;

[0013] A data processing module is used to generate search suggestion words based on the search characters and each character string, determine the weight corresponding to each character string as the weight of the corresponding search suggestion word; sort each search suggestion word according to the weight, and display the sorted search suggestion words through the search page of the terminal.

[0014] In one of the embodiments, the device also includes a dictionary tree construction module, which is used to obtain business data, wherein the business data includes business keywords and corresponding search popularity; determine the weight of each business keyword based on the search popularity corresponding to each business keyword; perform word segmentation processing on each business keyword to obtain a common prefix word of each business keyword; based on the common prefix word, generate a business keyword set including the common prefix word according to each business keyword; and construct a dictionary tree based on the business keywords in each business keyword set and the corresponding weights.

[0015] In one of the embodiments, the dictionary tree construction module is also used to determine the corresponding character string of each business keyword based on the common prefix word, and determine the character sequence of each business keyword in the business keyword set; construct at least one common prefix node of the dictionary tree based on the common prefix word and the character sequence of each business keyword, and construct corresponding child nodes according to the character sequence of each business keyword and the characters of the corresponding character string; store the weight of each business keyword in the character node corresponding to the corresponding character string.

[0016] In one of the embodiments, the nodes of the dictionary tree include a character space and a weight space, and the dictionary tree construction module is also used to store the characters corresponding to each string in the character space of the corresponding node; if the string includes a character subset, the weight corresponding to each business keyword is stored in the weight space of the last character node of the string; if the string does not include a character subset, a leaf node is configured at the end of the node of the string, an end character is configured for the leaf node, the end character is stored in the character space, and the weight corresponding to the string is stored in the weight space of the leaf node.

[0017] In one of the embodiments, the dictionary tree construction module is also used to obtain the first character of each business keyword and obtain the hash value corresponding to each first character; based on the hash value of each first character, determine the subtree cluster identifier corresponding to each business keyword; distribute the business keywords and the corresponding weights to the subtree cluster servers corresponding to each subtree cluster identifier, and construct a dictionary tree based on the received business keywords and the corresponding weights through each subtree cluster server.

[0018] In one of the embodiments, the dictionary tree construction module is also used to determine the target parameter value corresponding to each business keyword based on the hash value and a preset algorithm; based on the target parameter value, determine the storage location corresponding to each business keyword; and assign a corresponding subtree cluster identifier to each business keyword according to the storage location.

[0019] In one of the embodiments, the data search module is also used to search for corresponding character nodes in the dictionary tree according to the character sequence of the search character, determine the character node as a common prefix node, determine the last character node of the common prefix node as a direct prefix node, and determine a target search subtree according to the common prefix node and the direct prefix node; determine a search path in the target search subtree according to the direct prefix node, and traverse at least one child node corresponding to the direct prefix node according to the search path; obtain the character strings and weights corresponding to each traversed child node, and generate a character string set corresponding to the common prefix node based on each character string and the corresponding weight.

[0020] In one of the embodiments, the data search module is also used to traverse the child nodes corresponding to the direct prefix node in the target search subtree according to the search path; obtain the weight of the current character node, and when the weight of the current character node is empty, traverse the next node of the current character node; until the leaf node including the end character is traversed, the weight of the leaf node is obtained; determine the character string corresponding to the leaf node and the weight of the character string according to the search path.

[0021] In one of the embodiments, the data search module is also used to obtain the first character of the search characters and obtain a hash value corresponding to the first character; based on the hash value of the first character, determine the subtree cluster identifier corresponding to the search character; forward the search character to a target subtree cluster server corresponding to the subtree cluster identifier, and obtain corresponding search suggestion words and corresponding weights from a pre-built dictionary tree based on the search character through the target subtree cluster server.

[0022] In one of the embodiments, the data processing module is also used to determine the weight of the search keyword corresponding to each character string based on the weight of each character string; sort each search suggestion word in descending order based on the weight, and obtain the search suggestion word based on the sorting result and the quantity threshold; return the obtained search suggestion word to the terminal, and display the search suggestion word through the search page of the terminal according to the sorting result.

[0023] In one of the embodiments, the device also includes a dictionary tree update module, which is used to obtain historical business data according to a preset frequency, and the historical business data includes business keywords and corresponding current search popularity; determine the current weight of each business keyword according to the current search popularity of each business keyword; and update the dictionary tree based on each business keyword and the corresponding current weight.

[0024] In one of the embodiments, the dictionary tree update module is also used to extract the updated business keywords and the corresponding current weights if the business keywords include updated business keywords; determine the corresponding updated characters and character sequences based on the prefix characters of the updated business keywords; and update the updated characters and the corresponding current weights in the dictionary tree based on the character sequence.

[0025] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0026] Get the search character;

[0027] According to the character sequence of the search character, a corresponding character node is searched in a pre-constructed dictionary tree, a sub-node sequence corresponding to the character node is traversed, a character string corresponding to the sub-node sequence is extracted from the dictionary tree, and a weight corresponding to each character string is obtained from a node corresponding to the character string;

[0028] Generating search suggestion words according to the search characters and each character string, and determining the weight corresponding to each character string as the weight of the corresponding search suggestion word;

[0029] The search suggestion words are sorted according to the weights, and the sorted search suggestion words are displayed on a search page of the terminal.

[0030] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:

[0031] Get the search character;

[0032] According to the character sequence of the search character, a corresponding character node is searched in a pre-constructed dictionary tree, a sub-node sequence corresponding to the character node is traversed, a character string corresponding to the sub-node sequence is extracted from the dictionary tree, and a weight corresponding to each character string is obtained from a node corresponding to the character string;

[0033] Generating search suggestion words according to the search characters and each character string, and determining the weight corresponding to each character string as the weight of the corresponding search suggestion word;

[0034] The search suggestion words are sorted according to the weights, and the sorted search suggestion words are displayed on a search page of the terminal.

[0035] The above-mentioned dictionary-based search processing method, device, equipment and storage medium, after obtaining the search character, finds the corresponding character node in the dictionary tree according to the search character, and traverses the dictionary tree according to the character node, thereby obtaining the character string and weight corresponding to the sub-node sequence. Since the dictionary tree stores the character strings and corresponding weights of multiple business keywords, when matching in the dictionary tree, it is possible to quickly and effectively obtain the character string matching the search term and the corresponding weight. Then, search suggestion words are generated according to the search character and the character string, the search suggestion words are sorted according to the weight, and the sorted search suggestion words are returned to the terminal for display. By configuring storage space for characters and weights in the dictionary tree respectively, the dictionary tree can be used to quickly and effectively match the search suggestion words corresponding to the search word, and the weight corresponding to the search suggestion words can also be obtained. There is no need to additionally calculate the weight of the search suggestion words and then sort them, which effectively improves the reading efficiency and search processing efficiency of business data. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 An application environment diagram of a dictionary tree-based search processing method in one embodiment;

[0037] Figure 2 A flowchart of a search processing method based on a dictionary tree in one embodiment;

[0038] Figure 3 A flowchart of steps for constructing a dictionary tree in one embodiment;

[0039] Figure 4 A schematic diagram of a local structure of a dictionary tree in one embodiment;

[0040] Figure 5 A flowchart of the steps of constructing a dictionary tree in a specific embodiment;

[0041] Figure 6 A flowchart of the steps of constructing a dictionary tree in another embodiment;

[0042] Figure 7 A schematic diagram of performing a search process on a search character in one embodiment;

[0043] Figure 8 A schematic diagram of an interface for pushing results of search suggestion words in one embodiment;

[0044] Fig. 9 It is a flowchart of a search processing method based on a dictionary tree in a specific embodiment;

[0045] Fig.10 is a structural block diagram of a search processing device based on a dictionary tree in one embodiment;

[0046] Fig.11 is a structural block diagram of a search processing device based on a dictionary tree in another embodiment;

[0047] Fig.12 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0049] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0050] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. The solution provided in the embodiment of the present application involves artificial intelligence cloud computing, distributed storage, big data processing technology and other technologies. By processing a large amount of search data, it is possible to effectively realize intelligent recommendation of search terms.

[0051] The dictionary tree-based search processing method provided in the present application can be applied to a computer device. The computer device can be a terminal or a server. It is understandable that the content sharing processing method provided in the present application can be applied to a terminal, a server, or a system including a terminal and a server, and is implemented through the interaction between the terminal and the server.

[0052] In one embodiment, the computer device may be a server. The dictionary tree-based search processing method provided in this application may be applied to Figure 1 In the application environment shown, the application environment includes a system of terminals and servers. Among them, the terminal 102 communicates with the server 104 through the network. The server 104 obtains the search character input through the terminal 102, and the server 104 searches for the corresponding character node in the pre-built dictionary tree according to the search character, and traverses the dictionary tree according to the character node, thereby obtaining the character string and weight corresponding to the sub-node sequence. Quickly and effectively obtain the character string matching the search term and the corresponding weight. The server 104 then generates search suggestion words according to the search character and the character string, sorts the search suggestion words according to the weight, and returns the sorted search suggestion words to the terminal 102 for display. Among them, the terminal 102 can be but not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices, and the server 104 can be implemented with an independent server or a server cluster consisting of multiple servers.

[0053] In one embodiment, Figure 2 As shown, a search processing method based on a dictionary tree is provided, and the application to a computer device is used as an example for explanation. The computer device may specifically be a terminal or a server.

[0054] In this embodiment, the following steps are included:

[0055] S202, obtaining search characters.

[0056] The search character represents the word to be searched entered by the user in the input box of the terminal. For example, the search character can be a complete search word to be searched by the user, or an incomplete partial search word. The search character can include letters, Chinese characters, numbers, operation symbols, punctuation marks and other symbols.

[0057] A search engine system may be deployed in the terminal, and the search page of the terminal includes a search box. The search box refers to an interactive control in a data search system, which is used to extract corresponding accurate content from a large amount of information according to the search characters entered in the search box. Among them, the search box may be an input box, which is used to obtain and store content such as text information entered by the user through the keyboard or mouse. The user can enter the search characters to be searched through the search box on the terminal.

[0058] Specifically, when the user enters the content to be searched in the search box of the terminal, the computer device can obtain the search characters entered by the user in the search box of the terminal in real time through the data search system. For example, the user can enter "public provident fund" in the search box of the terminal, and "public", "province" and "gold" are the search characters entered by the terminal.

[0059] S204, searching for corresponding character nodes in a pre-constructed dictionary tree according to the character sequence of the search character, traversing the sub-node sequence corresponding to the character node, extracting the character string corresponding to the sub-node sequence from the dictionary tree, and obtaining the weight corresponding to each character string from the node corresponding to the character string.

[0060] Among them, the dictionary tree is also called a word search tree, prefix tree or key tree. It is a tree structure and a variant of a hash tree. The tree structure refers to a data structure in which there is a one-to-many tree relationship between data elements. It is an important type of nonlinear data structure. The tree structure can represent a hierarchical relationship. The dictionary tree can be used to store not only letters, but also other data such as numbers. A tree can include roots, subtrees and leaves. The root is equivalent to the root node in the tree structure, the subtree is equivalent to multiple child nodes in the tree structure, and the leaf is equivalent to the leaf node in the tree structure. Each subtree can also have its own subtree. In the tree structure, the root node has no predecessor node, and each of the remaining nodes has one and only one predecessor node. The leaf node has no subsequent node, and the number of subsequent nodes of each of the remaining nodes can be one or more. Among them, multiple means at least two.

[0061] The dictionary tree includes common prefixes, which can represent the common prefix characters included in multiple strings. The common prefix can be one character or multiple consecutive characters. Using the common prefixes of strings in the dictionary tree to query can reduce query time and reduce comparisons between strings, thereby improving query efficiency.

[0062] In the traditional dictionary tree, the root node does not contain characters, and each node except the root node contains only one character. From the root node to a certain node, the characters on the path are connected to form the string corresponding to the node. Therefore, in the traditional search processing method, according to the search character, only multiple corresponding search suggestion words can be matched in the dictionary tree, and it is necessary to further calculate or obtain the weight or search popularity of the search suggestion words, and then sort the multiple search suggestion words according to the weight or search popularity.

[0063] The trie includes a root node and multiple child nodes. The root node does not contain a character and may include a weight. Each node other than the root node can store a character and can also store a corresponding weight. That is, each node other than the root node includes a character space and a weight space. The character space is used to store the character of each node, and the weight space is used to store the weight of the string corresponding to the node. The characters included in all the child nodes of each node are different. Among them, the weight of the root node is empty or an initial value. For example, the weight of the root node can be "0".

[0064] The trie can be pre-analyzed using a large amount of sample data to obtain multiple business keywords, and the weight of each business keyword is calculated according to the corresponding search popularity. Then, the trie is constructed using the analyzed multiple business keywords and their corresponding weights. Therefore, the trie pre-constructed by the computer device includes the characters and weights corresponding to multiple business keywords.

[0065] Among them, the character sequence represents a sequence formed by several character runs. The search characters input by the terminal include the input order, and the character sequence can be the corresponding character sequence formed according to the input order of the search characters. The node sequence can represent the sequence of nodes passed through on the path from any node to the root node. The child node sequence can be the sequence of nodes passed through on the path of traversing the corresponding child nodes from the current node.

[0066] Specifically, after the computer device receives the search characters input by the terminal, it searches for the corresponding character node in the trie according to the search characters. When the character node includes multiple child nodes, the character node corresponding to the search characters can be a common prefix node. Among them, the last character node of the character node corresponding to the search characters can be determined as the direct prefix node. The direct prefix node can be the predecessor node in the tree structure. The common prefix node can include multiple character nodes or only one character node. The direct prefix node is only one node. The computer device then traverses the child node sequence corresponding to the direct prefix node according to the character node, extracts the string corresponding to the child node sequence from the trie, and obtains the weight corresponding to the string from the node corresponding to the string, so as to quickly obtain the string and weight corresponding to the search characters in the trie.

[0067] For example, if the search characters input by the terminal are "provident fund", the character sequence of the search characters can be: "gong" -> "ji" -> "jin". The computer device then searches for the character node corresponding to this character sequence in the trie, and takes the character nodes corresponding to "gong", "ji", and "jin" as the common prefix nodes, and determines "jin" as the current direct prefix node. If the search characters are "gong", then "gong" is the common prefix node and also the direct prefix node. The child nodes corresponding to this node can include multiple child nodes such as "ji", "ping", and "gong".

[0068] The computer device traverses each sub-node sequence corresponding to the direct prefix node in the dictionary tree, and obtains the character string and weight corresponding to each sub-node sequence. Among them, the character string is the character corresponding to each sub-node of the byte sequence concatenated in the order of the character sequence. The character string can include multiple characters or only one character. For example, the common prefix node of "Provident Fund" can include multiple character strings such as "loan", "withdrawal", "query", etc., and each character string stores a corresponding weight.

[0069] S206: Generate search suggestion words according to the search characters and each character string, and determine the weight corresponding to each character string as the weight of the corresponding search suggestion word.

[0070] The search suggestion word indicates that in the subsequent search process, recommended information matching the search characters input by the user is pushed to the user.

[0071] After the computer device matches multiple character strings corresponding to the search character in the dictionary tree, it generates multiple search suggestion words using the common prefix characters corresponding to the search character and each character string, and the weight of the character string matched in the dictionary tree is the weight of the corresponding search suggestion word. For example, taking the search character "Provident Fund" as an example, the common prefix character corresponding to the search character is "Provident Fund", and the corresponding character strings may include "loan", "withdrawal", and "query". The search suggestion words generated according to the search character and each character string include "Provident Fund Loan", "Provident Fund Withdrawal" and "Provident Fund Query", and the weight of each character string is determined as the weight of the corresponding search suggestion word, thereby being able to quickly and effectively obtain the search suggestion word corresponding to the search character and the corresponding weight.

[0072] S208: sorting the search suggestion words according to the weights, and displaying the sorted search suggestion words on a search page of the terminal.

[0073] After obtaining the search suggestion words corresponding to the search character and the corresponding weight, the computer device sorts the multiple search suggestion words according to the weight of each search suggestion word, wherein the sorting method can be to sort the multiple search suggestion words in descending order according to the weight, thereby obtaining sorted search suggestion words. The computer device returns the sorted multiple search suggestion words to the terminal and displays them on the screen of the terminal according to the sorting result.

[0074] In one of the embodiments, after the computer device sorts the multiple search suggestion words in descending order according to the weights, it can also extract a preset number of search suggestion words, return the extracted multiple search suggestion words to the terminal, and display them according to the sorting results.

[0075] After returning the obtained sorted search suggestion words to the terminal for display, the user can select the displayed multiple search suggestion words. The computer device obtains the selection operation triggered by the user through the terminal for the selected target search suggestion word, enters the target search suggestion word into the search box, and displays the target search suggestion word in the search box of the terminal. The computer device can further search for related information based on the target search suggestion word.

[0076] For example, based on the search character "provident fund" input by the terminal, the computer device can match the character strings corresponding to multiple sub-node sequences in the dictionary tree, such as "loan", "withdrawal", "query" and other character strings and corresponding weights. The computer device generates multiple search suggestion words and weights corresponding to "provident fund loan", "provident fund withdrawal", and "provident fund query" according to the search characters and character strings. The computer device pushes the multiple search suggestion words to the terminal after sorting them according to the weights and displays them according to the weight sorting, so as to provide them to the user for selection. After the user selects the target search suggestion word, the computer device displays the target search suggestion word in the search box and searches for the corresponding data resources according to the target search suggestion word.

[0077] In a traditional dictionary tree, the root node does not contain any characters, and each node except the root node contains only one character. From the root node to a certain node, the characters on the path are connected to form the string corresponding to the node. Therefore, in the traditional search processing method, the computer device can only match multiple corresponding search suggestion words in the dictionary tree according to the search characters, and it is necessary to further calculate or obtain the weight or search popularity of the search suggestion words, and then sort the multiple search suggestion words according to the weight or search popularity and return them to the terminal.

[0078] The above-mentioned search processing method based on the dictionary tree, after obtaining the search character, searches for the corresponding character node in the pre-built dictionary tree according to the search character, and traverses the dictionary tree according to the character node, thereby obtaining the character string and weight corresponding to the sub-node sequence. Since the dictionary tree stores the character strings and corresponding weights of multiple business keywords, when matching in the dictionary tree, the character string matching the search term and the corresponding weight can be directly obtained from the dictionary tree quickly and effectively. Then, the search suggestion words are generated according to the search character and the character string, the search suggestion words are sorted according to the weight, and the sorted search suggestion words are returned to the terminal for display. By configuring the storage space for characters and weights in the dictionary tree respectively, the search suggestion words corresponding to the search word can be matched quickly and effectively using the dictionary tree, and the weight corresponding to the search suggestion words can also be obtained. There is no need to calculate the weight of the search suggestion words and then sort them, which effectively improves the reading efficiency of the search data and the search processing efficiency.

[0079] In one embodiment, Figure 3 As shown, before obtaining the search character, a step of constructing a dictionary tree is also included, which specifically includes the following contents:

[0080] S302, obtaining business data, where the business data includes business keywords and corresponding search popularity.

[0081] S304: Determine the weight of each business keyword based on the search popularity corresponding to each business keyword.

[0082] S306, performing word segmentation processing on each business keyword to obtain a common prefix word of each business keyword.

[0083] S308 , based on the common prefix, generate a business keyword set including the common prefix according to each business keyword.

[0084] S310, constructing a dictionary tree based on the business keywords in each business keyword set and the corresponding weights.

[0085] Among them, business data can be historical search data and topic data, etc. Business data can be data in a local database or sample data obtained by a computer device from several third-party databases. Business data includes multiple business keywords and corresponding search popularity. Search popularity can indicate the search frequency of sample data containing business keywords within a certain period of time. Weight refers to the importance of a factor or indicator relative to a certain thing, which can indicate the frequency with which business keywords are searched. For example, the higher the search popularity of a business keyword, the greater its corresponding weight.

[0086] Before the computer device obtains the search character, it needs to use a large amount of sample data to build a dictionary tree in advance. Specifically, the computer device can pre-deploy the tree structure of the dictionary tree, wherein the computer device first creates the root node of the dictionary tree, the root node has no predecessor node, and each of the remaining nodes has one and only one predecessor node. Character space and weight space are allocated to each node in the dictionary tree, the character space is used to store the characters of each node, and the weight space is used to store the weight of the string corresponding to the node. The characters contained in all child nodes of each node are different. Among them, the character of the root node of the dictionary tree is configured to be empty, the weight of the root node is empty, or is an initial value, for example, the weight of the root node can be "0".

[0087] The computer device then analyzes a large amount of business data, extracts multiple business keywords from the business data, and obtains the search popularity of each business keyword. In another embodiment, the business data includes corresponding search popularity, and the business data may include multiple business keywords. By analyzing the search popularity of the business data, the search popularity of each business keyword in the business data can be obtained.

[0088] Among them, multiple business keywords may include the same common prefix, and the business keywords include the common prefix and the corresponding business words, that is, the business keywords include the corresponding common prefix and the corresponding character string. The common prefix may include multiple characters or only one character. The computer device may generate a business keyword set from multiple business keywords including the same common prefix. The business keyword set may also include the weight of each business keyword.

[0089] In the process of constructing the dictionary tree, a node corresponding to the string is generated and inserted in the dictionary tree structure. The computer device can first convert the string to be inserted into a character array and process each character. At the same time, when inserting a node, it can be determined whether the inserted string is a common prefix of a word inserted in the dictionary tree, or whether a word inserted in the dictionary tree is a common prefix of the word. If the current string includes the common prefix of other strings, or other strings include the common prefix of the current string, the characters with the same common prefix are reused. Otherwise, the word is newly created as a corresponding child node in the dictionary tree.

[0090] The computer device then constructs a dictionary tree using multiple business keyword sets and the weights of the business keywords. Specifically, the computer device can construct a corresponding node using each character of the business keyword, and each node can store the corresponding character and weight. Among them, multiple business keywords in the business keyword set include the same common prefix word, and in the dictionary tree, multiple child nodes corresponding to business matching words can be directly constructed under the character node of the common prefix word.

[0091] In this embodiment, by constructing a tree structure including characters and weights, a dictionary tree storing character strings of multiple business keywords and corresponding weights can be effectively constructed, so that when matching in the dictionary tree, the character strings matching the search terms and the corresponding weights can be quickly and effectively obtained.

[0092] In one embodiment, constructing a dictionary tree based on business keywords in each business keyword set and the corresponding weights includes: determining the character string corresponding to each business keyword based on a common prefix word, and determining the character sequence of each business keyword in the business keyword set; constructing at least one common prefix node of the dictionary tree based on the common prefix word and the character sequence of each business keyword, and constructing corresponding child nodes based on the character sequence of each business keyword and the characters of the corresponding character string; storing the weight of each business keyword in the node of the character string corresponding to the corresponding character string.

[0093] Among them, each node in the dictionary tree can store corresponding characters and weights. The business keyword set includes a common prefix word and a corresponding character string, and each business keyword consists of a common prefix word and a corresponding character string. The computer device can determine the character sequence of each business keyword in the business keyword set according to the order of the common prefix word and the corresponding character string. The character string corresponding to each business keyword in the business keyword set can be a character subset of the common prefix word. The common prefix word can include multiple characters or only one character.

[0094] In the process of constructing a dictionary tree using a set of business keywords, a computer device can construct a common prefix node based on a common prefix word and a corresponding character sequence. The common prefix node may include multiple character nodes or just one character node. Based on the prefix node, the computer device constructs multiple corresponding child nodes using the characters of multiple business words in the set according to the character sequence, with each character corresponding to a child node. When the computer device constructs a node in the dictionary tree, the corresponding character is stored in the node, and the weight of the business keyword is stored in the node of the character string corresponding to the business word.

[0095] The weight of the business keyword can be stored in the character subnode, or in the leaf node at the end of the string. Specifically, if the string includes a character subset, the weight of the string is stored in the weight space of the character node of the last character of the string. For example, the common prefix word "provident fund" can be followed by multiple business keywords such as "provident fund withdrawal", "provident fund inquiry", "provident fund loan", "provident fund storage", and "provident fund storage time". The character subset of "provident fund" can include "extraction", "inquiry", "loan", "storage", etc., among which the common prefix word "provident fund storage" can be followed by "provident fund storage time", and "provident fund storage" can also include the character subset "time". Therefore, when constructing the dictionary tree, the weight of "provident fund storage" can be stored in the subnode corresponding to the character "storage".

[0096] If the character string does not include a character subset, a leaf node is configured at the end of the node of the character string, and the computer device stores the weight of the business keyword in the leaf node corresponding to the character string.

[0097] In one of the embodiments, if the node also includes a child node, the weight of the node can be determined as a null value or an initial threshold. For example, the computer device can also determine the weight of the node containing the child node as "0". After the child node of the complete string of the business word is constructed, the corresponding leaf node is configured at the end of the child node, and the computer device can store a preset end character in the leaf node, and the end character is used to indicate that the character sequence of the corresponding business keyword ends here. The computer device also stores the weight of the business keyword of the sequence corresponding to the leaf node in the leaf node.

[0098] For example, Figure 4 As shown, take "Provident Fund Loan", "Provident Fund Withdrawal" and "Provident Fund Inquiry" as examples. Figure 4 It is a schematic diagram of the local structure of the dictionary tree. "Provident Fund Loan", "Provident Fund Withdrawal" and "Provident Fund Query" can be a business keyword set, and the business keyword set includes three business keywords. Among them, "Provident Fund Loan", "Provident Fund Withdrawal" and "Provident Fund Query" all contain the common prefix word of "Provident Fund", and "Loan", "Withdrawal" and "Query" can be the business words corresponding to the prefix word of "Provident Fund". The computer device can use "Provident Fund" as the common prefix word to construct the corresponding common prefix node, "Gold" as the direct prefix node, and construct multiple sub-nodes of the character string corresponding to the business words based on the direct prefix node of "Gold". For example Figure 4 As shown in , the first node R is the root node, and the multiple nodes n under the root node are all child nodes, and the terminal node l behind the child node n is a leaf node. Among them, n1, n2, and n3 are the common prefix nodes corresponding to "provident fund loan", "provident fund withdrawal", and "provident fund query", and n3 is a direct prefix node. The computer device then constructs a corresponding leaf node at the end of the child node of each string, and stores a preset end character in the leaf node. For example, the end character can be "$", indicating the end of the string of the node path, which can be used to identify the traversal end of the string. Among them, the weight of the non-leaf node can be a null value, for example, it can be "0". The tree structure constructed by the common prefix word of "provident fund" and the corresponding "loan", "withdrawal", "query" and other business words can be one of the subtrees in the dictionary tree. In this way, the computer device can effectively store the characters and weights corresponding to the business keywords.

[0099] In one embodiment, Figure 5 As shown, a specific step of constructing a dictionary tree is provided, including the following contents:

[0100] S502, obtaining business data, where the business data includes business keywords and corresponding search popularity.

[0101] S504: Determine the weight of each business keyword based on the search popularity corresponding to each business keyword.

[0102] S506: Perform word segmentation processing on each business keyword to obtain a common prefix word of each business keyword.

[0103] S508 , based on the common prefix, generate a business keyword set including the common prefix according to each business keyword.

[0104] S510, determining a character string corresponding to each business keyword according to the common prefix, and determining a character sequence of each business keyword in a business keyword set.

[0105] S512, construct at least one common prefix node of the dictionary tree based on the common prefix word and the character sequence of each business keyword, and construct corresponding child nodes according to the character sequence of each business keyword and the characters of the corresponding character string.

[0106] S514, storing the characters corresponding to each string into the character space of the corresponding node.

[0107] S516: If the character string includes a character subset, the weight corresponding to each business keyword is stored in the weight space of the last character node of the character string.

[0108] S518, if the string does not include a character subset, configure a leaf node at the end of the node of the string, configure an end character for the leaf node, store the end character in the character space, and store the weight corresponding to the string in the weight space of the leaf node.

[0109] Among them, when constructing a dictionary tree, the computer device can configure character space and weight space for each node of the dictionary tree. The character space is used to store the characters of each node, and the weight space is used to store the weight of the string corresponding to the node. Each node can deploy two parts of storage space, namely character space and weight space. The storage method of the weight space in the dictionary tree can be configured as incremental storage. For example, the character space can be a character type, which usually occupies 2 bytes of storage capacity. The weight space can be a long integer (Long type), which usually occupies 4 bytes of storage capacity.

[0110] The dictionary tree includes a root node and multiple child nodes. The root node does not contain characters, but includes weights. Each node except the root node can store characters and corresponding weights. That is, each node except the root node includes character space and weight space. All child nodes of each node contain different characters. Among them, the weight of the root node is empty or the initial value, for example, the weight of the root node can be "0".

[0111] In the process of constructing the dictionary tree, the computer device uses multiple business keywords to construct corresponding common prefix nodes and child nodes. The computer device stores the characters of each node in the corresponding character space, and the weight of the business keyword can be stored in the weight space of the child node or in the weight space of the leaf node. The end character is configured for the leaf node and the end character is stored in the character space.

[0112] Specifically, if the last character of the currently configured string also includes a character subset, the weight of the string is stored in the weight space of the character node of the last character of the string. If the last character of the currently configured string does not include a character subset, it means that the string is a complete string. At this time, a leaf node is configured at the end of the last character node of the string, and an end character is configured for the leaf node, and the end character is stored in the character space. The computer device also obtains the weight of the business keyword corresponding to the leaf node and stores the weight in the weight space of the leaf node. Thereby, the characters and weights corresponding to the business keywords can be effectively stored.

[0113] The dictionary tree in this embodiment requires more storage space than the traditional dictionary tree, but it stores the character strings and the corresponding weights at the same time. When searching for matching character strings in the dictionary tree, the corresponding weights are also obtained, thereby reducing the amount of computation and resource usage of additional weight calculations, effectively improving the efficiency of data search, thereby being able to quickly and effectively find matching search suggestion words.

[0114] In one embodiment, a dictionary tree is constructed based on business keywords in each business keyword set and the corresponding weights, including: obtaining the first character in each business keyword, obtaining the hash value corresponding to each first character; determining the subtree cluster identifier corresponding to each business keyword based on the hash value of each first character; distributing each business keyword and the corresponding weight to the subtree cluster server corresponding to each subtree cluster identifier, and constructing a dictionary tree based on the received business keywords and the corresponding weights through each subtree cluster server.

[0115] Among them, the hash value is a digital "fingerprint" created for data based on a hash function (or hash algorithm, also known as a hash function). The message or data is compressed into a summary through the hash function, which makes the data volume smaller and fixes the format of the data. The hash function scrambles the data and recreates a data fingerprint, namely the hash value. The hash value is usually represented by a short string of random letters and numbers. The hash function can map each character element to a series of conflict-free integer mapping values, and the resulting mapping value is the hash value.

[0116] After the computer device obtains each business keyword and the corresponding weight, it obtains the first character of each business keyword, that is, the first character of the business keyword. The computer device then obtains the hash value of the first character. Specifically, the hash value of each character can be pre-stored in the computer device. After the computer device obtains the first character of each business keyword, it directly searches the storage library for the hash value corresponding to the first character. In another embodiment, after the computer device obtains the first character of each business keyword, it can also calculate the corresponding hash value in real time according to a preset algorithm.

[0117] The computer device then determines the subtree cluster identifier corresponding to each business keyword based on the hash value of the first character in each business keyword, and distributes each business keyword and the corresponding weight to the subtree cluster server corresponding to each subtree cluster identifier. After obtaining the business keyword and the corresponding weight, each subtree cluster server performs word segmentation processing on each business keyword to obtain the common prefix of each business keyword. Based on the common prefix, a business keyword set including the common prefix is ​​generated according to each business keyword. The character string corresponding to each business keyword is determined according to the common prefix, and the character sequence of each business keyword in the business keyword set is determined; at least one common prefix node of the dictionary tree is constructed based on the common prefix and the character sequence of each business keyword, and the corresponding subnode is constructed according to the character sequence of each business keyword and the characters of the corresponding character string; the weight of each business keyword is stored in the node of the character string corresponding to the corresponding character string, thereby constructing the corresponding dictionary tree.

[0118] In one embodiment, the computer device may be a cluster server, which includes a routing server and at least one subtree cluster server. After the routing server obtains the business data including business keywords and the corresponding search popularity, it obtains the hash value of the first character in each business keyword, determines the subtree cluster identifier corresponding to each business keyword based on the hash value of the first character in each business keyword, and distributes each business keyword and the corresponding weight to the subtree cluster server corresponding to each subtree cluster identifier. Each subtree cluster server performs word segmentation on each business keyword to obtain a common prefix word of each business keyword; based on the common prefix word, a business keyword set including the common prefix word is generated according to each business keyword; and a corresponding dictionary tree is constructed based on the business keywords in each business keyword set and the corresponding weights.

[0119] In this embodiment, by constructing multiple distributed subtree clusters and multiple dictionary tree structures, each of which is a subset of a weighted dictionary tree, and each subset is deployed in a cluster, a large amount of data can be expanded and stored, thereby efficiently processing a large amount of data concurrently, thereby effectively improving the storage efficiency and processing efficiency of data.

[0120] In one embodiment, as Figure 6 shown, the steps of constructing the trie tree specifically include the following content:

[0121] S602, obtain the first characters in each service keyword, and obtain the hash values corresponding to the first characters.

[0122] S604, determine the target parameter values corresponding to each service keyword based on the hash values and a preset algorithm.

[0123] S606, determine the storage locations corresponding to each service keyword based on the target parameter values; allocate corresponding subtree cluster identifiers to each service keyword according to the storage locations.

[0124] S608, distribute each service keyword and its corresponding weight to the subtree cluster servers corresponding to the subtree cluster identifiers, and construct a trie tree through each subtree cluster server based on the received service keywords and their corresponding weights.

[0125] After the computer device obtains each service keyword and its corresponding weight, it obtains the first characters in each service keyword and obtains the hash value of the first character. The computer device further determines the target parameter values corresponding to each service keyword based on the hash values and a preset algorithm, and determines the storage locations corresponding to each service keyword based on the target parameter values.

[0126] Specifically, the computer device can calculate the hash value using a modulo algorithm or a consistent hashing algorithm to obtain the corresponding target parameter value. For example, the calculation formula of the hash value can be: s[0]*31^(n - 1)+s[1]*31^(n - 2)+...+s[n - 1]. Where s[0] is the ascii value of the first character of the pinyin string array, s[1] is the ascii value of the second character of the pinyin string array, and n is the length of the pinyin string.

[0127] For example, taking the service keyword "provident fund withdrawal" as an example, obtain the first character of "provident fund withdrawal", that is, "公", and the corresponding string can be "g", "o", "n", "g". Then the hash value corresponding to "公" can be calculated according to the preset algorithm. Thus, the hash value of "公" is s[0]*31^(n - 1)+s[1]*31^(n - 2)+...+s[n - 1] = 103*31^(4 - 1)+111*31^(4 - 2)+110*31^(4 - 3)+103 = 103*31*31*31+111*31*31+110*31+103 = 3178657.

[0128] In one embodiment, the computer device may use a modulus algorithm to take a modulus of a hash value to obtain a corresponding target parameter value. Specifically, the computer device takes a modulus of m on the hash value, and then the storage location of the string in the corresponding subtree may be determined based on the final target parameter value obtained. If the string is in Chinese, the string may be converted into corresponding pinyin characters. Specifically, the calculation formula may be as follows:

[0129] Target parameter value = hash value (pinyin string) % m, where m is a preset coefficient.

[0130] The computer device then modulos the hash value according to the above preset algorithm to obtain the corresponding target parameter value. For example, taking the business keyword "provident fund withdrawal" as an example, the hash value of the first character of "provident fund withdrawal" is 3178657, then the hash value is calculated to obtain 3178657%3=1, that is, the hash value is moduloed, and the corresponding remainder is 1, and the target parameter value is 1. Then, according to the obtained target parameter value, it can be determined that the business keyword "provident fund withdrawal" is stored in the corresponding second subtree, that is, according to the target parameter value, the storage location corresponding to each business keyword is determined. In this way, each business keyword can be effectively assigned a corresponding subtree cluster identifier, thereby effectively allocating each business keyword and the corresponding weight to the corresponding storage location, that is, the corresponding subtree cluster server.

[0131] In another embodiment, the computer device can use a consistent hashing algorithm to calculate the hash value to obtain the corresponding target parameter value. Specifically, the computer device dictionary tree cluster obtains the cluster hash value corresponding to each subtree cluster identifier, and configures the mapping space based on each cluster hash value; based on the hash value, the keyword hash value corresponding to each business keyword is determined according to the consistent hashing algorithm; based on the mapping space, the cluster hash value corresponding to each keyword hash value is determined, and the corresponding subtree cluster identifier is assigned to each business keyword according to the cluster hash value. Among them, the computer device can first map each subtree cluster server to a hash mapping space, for example, the hash mapping space can be a hash ring of a ring structure, and the cluster hash values ​​corresponding to each subtree cluster identifier are arranged in the hash mapping space with corresponding hash values. The computer device then maps the hash values ​​of each business keyword to the hash mapping space according to the consistent hashing algorithm. After each cluster hash value and each keyword hash value are mapped to the same hash ring, to determine which node a certain business keyword object is mapped to, it is only necessary to start from the object and search clockwise along the hash ring to find the first subtree cluster node, that is, the cluster hash value corresponding to the subtree cluster node is determined as the target parameter value corresponding to the business keyword. And the subtree cluster identifier corresponding to the business keyword is determined based on the target parameter value.

[0132] By adopting the consistent hashing algorithm, the subtree cluster identifier corresponding to each data is allocated based on the hash value, so that the data of the subtree cluster can be more balanced and the load balance of each subtree cluster can be guaranteed.

[0133] In one embodiment, in order to break the CPU and memory limitations of a single machine, a preset data segmentation rule can be configured to segment the data of the original single-course dictionary tree, and the concurrent support required by the business system, such as how many qps needs to be supported per second, can be used to perform a stress test on the segmented subtree. If it is not satisfied, it will continue to be split into more subtrees; in another embodiment, the number of subtree clusters can also be determined according to the number of virtual machines. Assuming that there are n virtual hosts available, the dictionary tree is segmented into n subtrees, and the business data is stored based on the corresponding n subtree clusters, thereby effectively ensuring the balance and effectiveness of data storage.

[0134] In one embodiment, a corresponding character node is searched in a pre-built dictionary tree according to a character sequence of a search character, a sub-node sequence corresponding to the character node is traversed, a character string corresponding to the sub-node sequence is extracted from the dictionary tree, and a weight corresponding to each character string is obtained from a node corresponding to the character string, including: searching for a corresponding character node in a dictionary tree according to a character sequence of a search character, determining the character node as a common prefix node, determining the last character node of the common prefix node as a direct prefix node, determining a target search subtree according to the common prefix node and the direct prefix node; determining a search path in the target search subtree according to the direct prefix node, and traversing at least one sub-node corresponding to the direct prefix node according to the search path; obtaining a character string and a weight corresponding to each traversed sub-node, and generating a character string set corresponding to the common prefix node based on each character string and the corresponding weight.

[0135] The computer device may use the last node of the direct prefix node as the search root node of the target search subtree, and the search path may be a path from the search root node to each leaf node.

[0136] After the computer device receives the search character input by the terminal, it searches for the corresponding character in the dictionary tree according to the character sequence of the search character and matches the corresponding character node. The character node is determined as a common prefix node, which can be one or more. When the search character is one character, the character node corresponding to the character is matched in the dictionary tree, and the corresponding character node is determined as a common prefix node and a subtree root node, and the multiple subnodes corresponding to the common prefix node are used as the target search subtree. If the search character includes multiple characters, the computer device first searches for the first character input by the terminal in the dictionary tree, and continues to search for the second character input by the terminal in the dictionary tree according to the character sequence until all the character nodes corresponding to the search character are found, and the multiple character nodes corresponding to the search character are determined as common prefix nodes, and the last character node of the common prefix node is determined as a direct prefix node, and the target search subtree is determined according to the common prefix node and the direct prefix node. The computer device then uses the common prefix node, the direct prefix node and the corresponding several subnodes as the target search subtree.

[0137] After the computer device determines the common prefix node corresponding to the search character and the target search subtree to be searched, it searches for the corresponding child node in the target search subtree. Specifically, the computer device determines the search path in the target search subtree based on the direct prefix node, and searches for the corresponding multiple child nodes according to the search path. Each child node includes a corresponding character, and the leaf node includes the weight of the string corresponding to the search path. The computer device generates a string set using the strings of the child nodes traversed by each search path. The string set also includes the weight of each string. The common prefix node in the target search subtree is the common prefix character corresponding to the string set.

[0138] The computer device may further generate a plurality of search suggestion words respectively using the common prefix character and a plurality of character strings in the character string set, and the weight of each search suggestion word is the weight of the corresponding character string.

[0139] In one embodiment, if the search character does not have a corresponding character node in the dictionary tree, it means that the search word does not exist in the dictionary tree, and the computer device does not return the search suggestion word to the terminal. If the corresponding character node is found in the dictionary tree according to the search character, but there is no child node under the character node except the leaf node, it means that the search character does not have a search suggestion word in the dictionary tree, and the computer device does not return the search suggestion word to the terminal.

[0140] In this embodiment, by performing a prefix query in the dictionary tree according to the search character, the character string corresponding to the search character can be quickly and effectively found in the dictionary tree, and the weight corresponding to the character string can also be quickly and effectively obtained.

[0141] In one embodiment, the character string and weight corresponding to each traversed child node are obtained, and a character string set corresponding to a common prefix node is generated based on each character string and the corresponding weight, including: traversing the child nodes corresponding to the direct prefix node in the target search subtree according to the search path; obtaining the weight of the current character node, and when the weight of the current character node is empty, traversing the next node of the current character node; until the leaf node including the end character is traversed, the weight of the leaf node is obtained; and determining the character string corresponding to the leaf node and the weight of the string according to the search path.

[0142] Each node in the dictionary tree stores the corresponding character and weight, and the leaf node of the dictionary tree stores the end character and the weight of the string of the corresponding path of the leaf node.

[0143] After the computer device obtains the search character, it searches for the corresponding character in the dictionary tree according to the character sequence of the search character, matches the corresponding character node, determines the corresponding character node as the common prefix node, and determines the direct prefix node as the search root node, which can be a subtree root node. The computer device also uses multiple child nodes corresponding to the direct prefix node as the target search subtree. The computer device traverses the target search subtree according to the search path.

[0144] Specifically, the computer device determines the search path in the target search subtree according to the common prefix node and the direct prefix node. The computer device may use the entire common prefix node as the search root node of the target search subtree, or may use the direct prefix node as the search root node of the target search subtree. The computer device then searches for the corresponding multiple child nodes respectively starting from the search root node of the target search subtree according to the search path.

[0145] Each child node in the dictionary tree includes a corresponding character and a weight. If the character of the child node is not a preset end character, it means that the node still has child nodes and is not a terminal node, and the weight of the node is empty, for example, it can also be a preset threshold representing an empty value. In the process of traversing the child nodes corresponding to the direct prefix node, the computer device obtains the weight of the current character node traversed. When the weight of the character node is empty, it means that the character node also includes a corresponding child node, and the computer device continues to traverse the next node of the character node until it traverses to a leaf node with an end character, indicating that the traversal of the search path is completed. The computer device obtains the weight stored in the leaf node, and the computer device can then effectively determine the character string corresponding to the leaf node under the search path and the weight of the character string based on the search path.

[0146] In this embodiment, a string set is generated using the string of the child node traversed by each search path, so the string set includes multiple strings corresponding to the search character and the weight of each string. Thus, the computer device can quickly and effectively search for the string corresponding to the search character in the dictionary tree, and can also quickly and effectively obtain the weight corresponding to the string.

[0147] In one embodiment, a corresponding character node is searched in a pre-built dictionary tree according to a character sequence of a search character, a sub-node sequence corresponding to the character node is traversed, a character string corresponding to the sub-node sequence is extracted from the dictionary tree, and a weight corresponding to each character string is obtained from a node corresponding to the character string, including: obtaining the first character in the search character, and obtaining a hash value corresponding to the first character; based on the hash value of the first character, determining a sub-tree cluster identifier corresponding to the search character; forwarding the search character to a target sub-tree cluster server corresponding to the sub-tree cluster identifier, and obtaining a corresponding search suggestion word and a corresponding weight from a pre-built dictionary tree based on the search character through the target sub-tree cluster server.

[0148] Among them, multiple subtree clusters can be deployed in the computer device by constructing multiple dictionary tree structures, each of which is a subset of the weighted dictionary tree, and each subset is deployed in a cluster manner.

[0149] After searching for a character, the computer device obtains the hash value corresponding to the first character of the string corresponding to the search character; based on the hash value, the storage location of the string in the corresponding subtree, that is, the corresponding subtree cluster identifier, is determined. The computer device then forwards the search character to the subtree cluster server corresponding to the subtree cluster identifier; the subtree cluster server searches for the corresponding search suggestion words and their corresponding weights in the pre-built dictionary tree according to the search character, and returns them to the computer device. The computer device sorts the search suggestion words according to the weights and returns them to the terminal. By storing the search data in a distributed subtree cluster, when the amount of data is large, the search data can be processed quickly and efficiently, and the corresponding search suggestion words can be efficiently obtained from the corresponding subtree cluster, thereby effectively improving the concurrent processing efficiency and data reading efficiency of the data.

[0150] In one of the embodiments, the dictionary tree cluster includes a routing server and multiple subtree cluster servers, and each subtree cluster server stores a subtree corresponding to the corresponding search data. The user can enter the search character through the search page of the terminal, and the terminal initiates a search request based on the search character, and the terminal can send the search request to the routing server through the gateway. After the routing server obtains the search request and the search character carried in the search request, it obtains the hash value corresponding to the first character in the search character; according to the hash value, it determines the subtree cluster identifier corresponding to the storage location of the search string. The routing server then forwards the search request to the subtree cluster server corresponding to the subtree cluster identifier; the subtree cluster server then searches for the corresponding search suggestion words and the corresponding weights in the pre-built dictionary tree according to the search character, and returns them to the routing server. The routing server sorts the search suggestion words according to the weights and returns them to the terminal, and displays the search suggestion words on the search page of the terminal. For example, Figure 7 As shown, it is a schematic diagram of searching for search characters in one embodiment, and the dictionary tree cluster may include multiple dictionary tree subtree clusters with weights.

[0151] In this embodiment, in a distributed dictionary tree cluster, when a user inputs a prefix string and initiates a search request through a terminal, the routing server determines which subtree cluster server is to process the request based on the hash value of the first character of the search string, and forwards the search request to the corresponding subtree cluster server, so that the corresponding subtree cluster server processes the search request respectively. When the amount of data is large, the search data can be processed quickly and efficiently, and the corresponding search suggestion words can be efficiently obtained from the corresponding subtree cluster, thereby effectively improving the concurrent processing efficiency and data reading efficiency of the data.

[0152] In one embodiment, each search suggestion word is sorted according to a weight, and the sorted search suggestion words are returned to the terminal for display, including: determining the weight of the search keyword corresponding to each character string based on the weight of each character string; sorting each search suggestion word in descending order based on the weight, and obtaining the search suggestion words based on the sorting result and the quantity threshold; returning the obtained search suggestion words to the terminal, and displaying the search suggestion words through the search page of the terminal according to the sorting result.

[0153] After the computer device finds a string set matching the search character in the dictionary tree according to the search character, it can generate multiple search suggestion words respectively using the search character and multiple strings in the string set. Specifically, the computer device can use the search character as a common prefix character, and use the common prefix character to combine with each string in the string set, so as to obtain a corresponding search suggestion word respectively.

[0154] For example, taking the search character as "provident fund" as an example, the computer device searches for the corresponding character nodes in the trie based on "provident fund", and uses these three character nodes as common prefix nodes to traverse and search for several child nodes corresponding to the common prefix nodes in the trie. Refer to Figure 4 , the common prefix nodes corresponding to "provident fund" may include multiple child nodes such as "loan", "withdrawal", and "inquiry". The computer device then obtains the strings and weights corresponding to the multiple traversed child nodes, and generates a set of strings corresponding to the common prefix nodes using the multiple strings and weights. Therefore, the generated set of strings may include multiple strings such as "loan / 24", "withdrawal / 45", "inquiry / 21" and the corresponding weights. The computer device then generates multiple search suggestion words respectively using the search character and the multiple strings in the set of strings, so as to obtain search suggestion words such as "provident fund loan (weight 24)", "provident fund withdrawal (weight 45)", "provident fund inquiry (weight 21)" and the corresponding weights. The computer device performs a descending sort according to the weights, that is, sorts from high to low according to the weights. Among them, if the number of search suggestion words exceeds the preset number threshold, extract the search suggestion words with the preset number threshold from the sorted multiple search suggestion words, and return the extracted search suggestion words to the terminal according to the sorting result and display them on the search interface of the terminal. Refer to Figure 8 shown, Figure 8 is a schematic diagram of the push result interface for returning corresponding search suggestion words according to the search character after the user inputs the search character "provident fund" through the terminal.

[0155] In one embodiment, the above search processing method based on the trie further includes: obtaining historical business data according to a preset frequency, where the historical business data includes business keywords and the corresponding current search popularity; determining the current weights of the business keywords according to the current search popularity of each business keyword; and updating the trie based on each business keyword and the corresponding current weight.

[0156] Among them, the storage method of the weight of each node in the trie is incremental storage. Since the initially constructed trie stores the weights of multiple characters and multiple strings, and during the operation of the data search system, the search popularity of some strings may change, and the corresponding weights will also change accordingly. Therefore, the computer device can update the trie using the historical business data within a period of time according to the preset frequency to update the corresponding characters and weights in the trie to ensure the accuracy of the weights of the business keywords stored in the trie.

[0157] Specifically, the computer device obtains some historical business data from a local database or a third-party database according to a preset frequency. The historical business data can be business data, search data, and topic data in the past period of time. The historical business data includes corresponding business keywords and current search popularity. The computer device extracts keywords from the historical business data to obtain multiple business keywords and corresponding current search popularity, and calculates the current weight of the corresponding business keyword according to the current search popularity.

[0158] The computer device searches for the character string and the weight of the character string corresponding to the business keyword in the dictionary tree, compares the weight of the character string in the dictionary tree with the current weight of the business keyword, and if the current weight is different from the weight stored in the dictionary tree, updates the weight of the character string in the dictionary tree to the current weight, thereby updating the weight of the corresponding business keyword in the dictionary tree. In this way, the business keywords and weights in the dictionary tree can be dynamically updated according to historical business data, thereby effectively ensuring the validity and accuracy of the business keywords and corresponding weights stored in the dictionary tree.

[0159] In one embodiment, a dictionary tree is updated based on each business keyword and the corresponding current weight, including: if the business keywords include updated business keywords, extracting the updated business keywords and the corresponding current weights; determining the corresponding updated characters and character sequences based on the prefix characters of the updated business keywords; and updating the updated characters and the corresponding current weights in the dictionary tree based on the character sequences.

[0160] After the computer device obtains a number of historical business data according to a preset frequency, it extracts keywords from the historical business data to obtain multiple business keywords and corresponding current search popularity, and calculates the current weight of the corresponding business keyword according to the current search popularity. The computer device further compares the extracted business keywords with the existing business keywords in the database. The database can be a local database or a third-party database. If the extracted business keywords include business keywords that are not in the database, it means that the business keyword is a newly added business keyword, and the business keyword is determined as an updated business keyword.

[0161] The computer device needs to update the updated service keywords and the corresponding current weights into the dictionary tree. Specifically, the computer device identifies whether there are common prefix characters of the updated service keywords in the dictionary tree. If there are partial prefix characters of the updated service keywords in the dictionary tree, the computer device determines the corresponding updated characters and character sequences based on the common prefix characters of the service keywords, and splits the updated service keywords into common prefix characters and updated characters. The computer device adds corresponding updated character nodes to the dictionary tree according to the character sequence, adds corresponding leaf nodes at the end of the updated character nodes, stores preset end characters in the leaf nodes, and stores the current weight of the updated service keywords in the leaf nodes.

[0162] If the common prefix character of the updated business keyword does not exist in the dictionary tree, the computer device determines all characters of the business keyword as updated characters, and adds an updated character node corresponding to the updated character in the dictionary tree according to the character sequence of the business keyword, and adds the current weight of the updated business keyword to the corresponding leaf node. In this way, the updated business keyword and the corresponding current weight can be effectively updated to the dictionary tree. In this way, the business keywords and weights in the dictionary tree can be dynamically updated according to historical business data, thereby effectively ensuring the validity and accuracy of the business keywords and corresponding weights stored in the dictionary tree.

[0163] In a specific embodiment, Fig. 9 As shown, a specific search processing method based on a dictionary tree is provided, comprising the following steps:

[0164] S902, obtaining search characters.

[0165] S904, search for corresponding character nodes in the dictionary tree according to the character sequence of the search character, determine the character nodes as common prefix nodes, determine the last character node of the common prefix nodes as direct prefix nodes, and determine the target search subtree according to the common prefix nodes and direct prefix nodes.

[0166] S906: Determine a search path in the target search subtree according to the direct prefix node, and traverse at least one child node corresponding to the direct prefix node according to the search path.

[0167] S908 , traverse the child nodes corresponding to the direct prefix node in the target search subtree according to the search path.

[0168] S910, obtaining the weight of the current character node, and when the weight of the current character node is empty, traversing the next node of the current character node.

[0169] S912, until the leaf node including the end character is traversed, the weight of the leaf node is obtained; and the character string corresponding to the leaf node and the weight of the character string are determined according to the search path.

[0170] S914: Determine the weight of the search keyword corresponding to each character string based on the weight of each character string.

[0171] S916 , sorting the search suggestion words in descending order based on the weights, and acquiring the search suggestion words based on the sorting result and the quantity threshold.

[0172] S918: Return the acquired search suggestion words to the terminal, and display the search suggestion words through the search page of the terminal according to the sorting result.

[0173] In this embodiment, since the dictionary tree stores the character strings and corresponding weights of multiple business keywords, when matching in the dictionary tree, the character strings matching the search terms and the corresponding weights can be quickly and effectively obtained directly from the dictionary tree. Then, search suggestion words are generated based on the search characters and character strings, the search suggestion words are sorted according to the weights, and the sorted search suggestion words are returned to the terminal for display. By configuring storage space for characters and weights in the dictionary tree respectively, the dictionary tree can be used to quickly and effectively match the search suggestion words corresponding to the search terms, and the weights corresponding to the search suggestion words can also be obtained. There is no need to additionally calculate the weights of the search suggestion words and then sort them, which effectively improves the efficiency of the search process.

[0174] The present application also provides an application scenario, which applies the above-mentioned dictionary tree-based search processing method. Specifically, the application of the dictionary tree-based search processing method in the application scenario is as follows:

[0175] A search application is deployed in the user terminal, and the search application includes a search engine system. The search page of the terminal includes a search box. The user can enter text information and other content through the keyboard or mouse. The terminal can enter the search character to be searched through the search box of the search page, and send a search request to the server corresponding to the search engine. After receiving the search request, the server obtains the search character carried by the search request, searches for the corresponding character node in the pre-built dictionary tree according to the search character, and traverses the dictionary tree according to the character node, thereby obtaining the character string and weight corresponding to the child node sequence. Since the dictionary tree stores the character strings and corresponding weights of multiple business keywords, when matching in the dictionary tree, the character string matching the search word and the corresponding weight can be directly obtained from the dictionary tree quickly and effectively. Then, the search suggestion word is generated according to the search character and the character string, the search suggestion word is sorted according to the weight, and the sorted search suggestion word is returned to the terminal for display, so that the corresponding search data can be obtained quickly and accurately, effectively improving the reading efficiency of the search data and the search processing efficiency.

[0176] It should be understood that although Figure 2 , 3 The steps in the flowcharts of , 5, 6, and 9 are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2 , 3 At least part of the steps in , 5, 6, and 9 may include multiple steps or multiple stages. These steps or stages do not necessarily have to be performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily one by one, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.

[0177] In one embodiment, Fig.10 As shown, a search processing device 1000 based on a dictionary tree is provided. The device can adopt a software module or a hardware module, or a combination of the two to become a part of a computer device. The device specifically includes: a data acquisition module 1002, a data search module 1004 and a data processing module 1006, wherein:

[0178] The data acquisition module 1002 is used to acquire the search character;

[0179] The data search module 1004 is used to search for corresponding character nodes in the pre-built dictionary tree according to the character sequence of the search character, traverse the sub-node sequence corresponding to the character node, extract the character string corresponding to the sub-node sequence from the dictionary tree, and obtain the weight corresponding to each character string from the node corresponding to the character string;

[0180] The data processing module 1006 is used to generate search suggestion words based on the search characters and each character string, determine the weight corresponding to each character string as the weight of the corresponding search suggestion word; sort each search suggestion word according to the weight, and display the sorted search suggestion words through the search page of the terminal.

[0181] In one embodiment, Fig.11As shown, the above-mentioned dictionary tree-based search processing device 1000 also includes a dictionary tree construction module 1001, which is used to obtain business data, the business data including business keywords and corresponding search popularity; determine the weight of each business keyword based on the search popularity corresponding to each business keyword; perform word segmentation processing on each business keyword to obtain a common prefix word of each business keyword; based on the common prefix word, generate a business keyword set including the common prefix word according to each business keyword; and construct a dictionary tree based on the business keywords in each business keyword set and the corresponding weights.

[0182] In one embodiment, the dictionary tree construction module 1001 is also used to determine the corresponding character string of each business keyword based on the common prefix word, and determine the character sequence of each business keyword in the business keyword set; construct at least one common prefix node of the dictionary tree based on the common prefix word and the character sequence of each business keyword, and construct corresponding child nodes according to the character sequence of each business keyword and the characters of the corresponding character string; store the weight of each business keyword in the character node corresponding to the corresponding character string.

[0183] In one embodiment, the nodes of the dictionary tree include character space and weight space, and the dictionary tree construction module 1001 is also used to store the characters corresponding to each string in the character space of the corresponding node; if the string includes a character subset, the weight corresponding to each business keyword is stored in the weight space of the last character node of the string; if the string does not include a character subset, a leaf node is configured at the end of the node of the string, an end character is configured for the leaf node, the end character is stored in the character space, and the weight corresponding to the string is stored in the weight space of the leaf node.

[0184] In one embodiment, the dictionary tree construction module 1001 is also used to obtain the first character of each business keyword and obtain the hash value corresponding to each first character; based on the hash value of each first character, determine the subtree cluster identifier corresponding to each business keyword; distribute each business keyword and the corresponding weight to the subtree cluster server corresponding to each subtree cluster identifier, and construct a dictionary tree based on the received business keywords and the corresponding weights through each subtree cluster server.

[0185] In one embodiment, the dictionary tree construction module 1001 is also used to determine the target parameter value corresponding to each business keyword based on the hash value and the preset algorithm; based on the target parameter value, determine the storage location corresponding to each business keyword; and assign a corresponding subtree cluster identifier to each business keyword according to the storage location.

[0186] In one embodiment, the data search module 1004 is also used to search for corresponding character nodes in the dictionary tree according to the character sequence of the search character, determine the character node as a common prefix node, determine the last character node of the common prefix node as a direct prefix node, and determine a target search subtree according to the common prefix node and the direct prefix node; determine a search path in the target search subtree according to the direct prefix node, and traverse at least one child node corresponding to the direct prefix node according to the search path; obtain the character string and weight corresponding to each traversed child node, and generate a character string set corresponding to the common prefix node based on each character string and the corresponding weight.

[0187] In one embodiment, the data search module 1004 is also used to traverse the child nodes corresponding to the direct prefix node in the target search subtree according to the search path; obtain the weight of the current character node, and when the weight of the current character node is empty, traverse the next node of the current character node; until the leaf node including the end character is traversed, the weight of the leaf node is obtained; determine the character string corresponding to the leaf node and the weight of the character string according to the search path.

[0188] In one embodiment, the data search module 1004 is also used to obtain the first character in the search characters and obtain the hash value corresponding to the first character; based on the hash value of the first character, determine the subtree cluster identifier corresponding to the search characters; forward the search characters to the target subtree cluster server corresponding to the subtree cluster identifier, and obtain the corresponding search suggestion words and corresponding weights from the pre-built dictionary tree based on the search characters through the target subtree cluster server.

[0189] In one embodiment, the data processing module 1006 is also used to determine the weight of the search keyword corresponding to each character string based on the weight of each character string; sort each search suggestion word in descending order based on the weight, and obtain the search suggestion word based on the sorting result and the quantity threshold; return the obtained search suggestion word to the terminal, and display the search suggestion word through the search page of the terminal according to the sorting result.

[0190] In one embodiment, the above-mentioned dictionary tree-based search processing device also includes a dictionary tree update module, which is used to obtain historical business data according to a preset frequency, and the historical business data includes business keywords and corresponding current search popularity; determine the current weight of each business keyword according to the current search popularity of each business keyword; and update the dictionary tree based on each business keyword and the corresponding current weight.

[0191] In one embodiment, the dictionary tree update module is also used to extract the updated business keywords and the corresponding current weights if the business keywords include updated business keywords; determine the corresponding updated characters and character sequences based on the prefix characters of the updated business keywords; and update the updated characters and the corresponding current weights in the dictionary tree based on the character sequence.

[0192] For the specific definition of the dictionary tree-based search processing device, please refer to the definition of the dictionary tree-based search processing method above, which will not be repeated here. Each module in the above-mentioned dictionary tree-based search processing device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0193] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig.12 As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store search characters, business keywords, dictionary trees, search suggestion words and other data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a search processing method based on a dictionary tree is implemented.

[0194] Those skilled in the art will understand that Fig.12 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0195] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0196] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0197] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0198] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0199] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A search processing method based on a dictionary tree, characterized in that: The method comprises: Get the search character; According to the character sequence of the search character, a corresponding character node is searched in a pre-constructed dictionary tree, the character node is determined as a common prefix node, the last character node of the common prefix node is determined as a direct prefix node, and a target search subtree is determined according to the common prefix node and the direct prefix node; Traversing at least one child node corresponding to the direct prefix node in the target search subtree; Get the string and weight corresponding to each traversed child node; Generating search suggestion words according to the search characters and each character string, and determining the weight corresponding to each character string as the weight of the corresponding search suggestion word; The search suggestion words are sorted according to the weights, and the sorted search suggestion words are displayed on a search page of the terminal.

2. The method according to claim 1, characterized in that Before obtaining the search character, the method further includes: Acquire business data, wherein the business data includes business keywords and corresponding search popularity; Determine the weight of each business keyword based on the search popularity corresponding to each business keyword; Perform word segmentation on each business keyword to obtain the common prefix of each business keyword; Based on the common prefix word, generating a business keyword set including the common prefix word according to each business keyword; A dictionary tree is constructed based on the business keywords in each business keyword set and the corresponding weights.

3. The method according to claim 2, characterized in that The constructing of a dictionary tree based on the business keywords in the business keyword set and the corresponding weights includes: Determine the character string corresponding to each business keyword according to the common prefix, and determine the character sequence of each business keyword in the business keyword set; Construct at least one common prefix node of the dictionary tree based on the common prefix word and the character sequence of each business keyword, and construct corresponding child nodes according to the character sequence of each business keyword and the characters of the corresponding character string; The weight of each business keyword is stored in the character node corresponding to the corresponding string.

4. The method according to claim 3, characterized in that The nodes of the dictionary tree include a character space and a weight space, and the weight of each business keyword is stored in the character node corresponding to the corresponding character string, including: Store the characters corresponding to each string into the character space of the corresponding node; If the character string includes a character subset, the weight corresponding to each business keyword is stored in the weight space of the last character node of the character string; If the string does not include a character subset, a leaf node is configured at the end of the node of the string, an end character is configured for the leaf node, the end character is stored in the character space, and the weight corresponding to the string is stored in the weight space of the leaf node.

5. The method according to any one of claims 2 to 4, characterized in that: The constructing a dictionary tree based on the business keywords in the business keyword set and the corresponding weights includes: Get the first character of each business keyword and get the hash value corresponding to each first character; Determine the subtree cluster identifier corresponding to each business keyword based on the hash value of each first character; The business keywords and the corresponding weights are distributed to the subtree cluster servers corresponding to the subtree cluster identifiers, and the subtree cluster servers construct a dictionary tree based on the received business keywords and the corresponding weights.

6. The method according to claim 5, characterized in that The determining of the subtree cluster identifier corresponding to each business keyword based on the hash value of each first character includes: Determine the target parameter value corresponding to each business keyword based on the hash value and a preset algorithm; Based on the target parameter value, determining the storage location corresponding to each business keyword; A corresponding subtree cluster identifier is allocated to each business keyword according to the storage location.

7. The method according to claim 1, characterized in that The traversing at least one child node corresponding to the direct prefix node in the target search subtree includes: A search path in the target search subtree is determined according to the direct prefix node, and at least one child node corresponding to the direct prefix node is traversed according to the search path.

8. The method according to claim 7, characterized in that The obtaining of the character string and weight corresponding to each traversed child node includes: Traversing the child nodes corresponding to the direct prefix node in the target search subtree according to the search path; Obtain the weight of the current character node, and when the weight of the current character node is empty, traverse the next node of the current character node; Until the leaf node including the end character is traversed, the weight of the leaf node is obtained; The character string corresponding to the leaf node and the weight of the character string are determined according to the search path.

9. The method according to any one of claims 7 to 8, characterized in that The search suggestion word and the corresponding weight are obtained by obtaining the first character of the search character and obtaining the hash value corresponding to the first character; Determine the subtree cluster identifier corresponding to the search character based on the hash value of the first character; The search character is forwarded to a target subtree cluster server corresponding to the subtree cluster identifier, and is obtained from a pre-constructed dictionary tree based on the search character by the target subtree cluster server.

10. The method according to claim 1, characterized in that The sorting of the search suggestion words according to the weights and returning the sorted search suggestion words to the terminal for display includes: Determine the weight of the search keyword corresponding to each character string based on the weight of each character string; Sort the search suggestion words in descending order based on the weights, and obtain the search suggestion words based on the sorting result and the quantity threshold; The acquired search suggestion words are returned to the terminal, and the search suggestion words are displayed through a search page of the terminal according to the ranking result.

11. The method according to claim 1, characterized in that: The method further comprises: Acquire historical business data according to a preset frequency, wherein the historical business data includes business keywords and corresponding current search popularity; Determine the current weight of each business keyword based on the current search popularity of each business keyword; The dictionary tree is updated based on each business keyword and the corresponding current weight.

12. The method according to claim 11, characterized in that The updating of the dictionary tree based on each business keyword and the corresponding current weight includes: If the service keywords include update service keywords, extract the update service keywords and the corresponding current weights; Determine the corresponding update character and character sequence according to the prefix character of the update service keyword; The updated characters and the corresponding current weights are updated in the dictionary tree based on the character sequence.

13. A search processing device based on a dictionary tree, characterized in that: The device comprises: A data acquisition module, used to acquire search characters; A data search module, used to search for corresponding character nodes in a pre-constructed dictionary tree according to the character sequence of the search character, determine the character node as a common prefix node, determine the last character node of the common prefix node as a direct prefix node, determine a target search subtree according to the common prefix node and the direct prefix node; traverse at least one child node corresponding to the direct prefix node in the target search subtree; and obtain the character string and weight corresponding to each traversed child node; A data processing module is used to generate search suggestion words based on the search characters and each character string, determine the weight corresponding to each character string as the weight of the corresponding search suggestion word; sort each search suggestion word according to the weight, and display the sorted search suggestion words through the search page of the terminal.

14. The device according to claim 13, characterized in that Also includes: A dictionary tree construction module is used to obtain business data, wherein the business data includes business keywords and corresponding search popularity; determine the weight of each business keyword based on the search popularity corresponding to each business keyword; perform word segmentation processing on each business keyword to obtain a common prefix word of each business keyword; Based on the common prefix, a business keyword set including the common prefix is ​​generated according to each business keyword; and a dictionary tree is constructed based on the business keywords in each business keyword set and the corresponding weights.

15. The device according to claim 14, characterized in that The dictionary tree construction module is also used to determine the corresponding character string of each business keyword based on the common prefix word, and determine the character sequence of each business keyword in the business keyword set; construct at least one common prefix node of the dictionary tree based on the common prefix word and the character sequence of each business keyword, and construct corresponding child nodes according to the character sequence of each business keyword and the characters of the corresponding character string; store the weight of each business keyword in the character node corresponding to the corresponding character string.

16. The device according to claim 15, characterized in that The nodes of the dictionary tree include a character space and a weight space, and the dictionary tree construction module is also used to store the characters corresponding to each string in the character space of the corresponding node; If the character string includes a character subset, the weight corresponding to each business keyword is stored in the weight space of the last character node of the character string; If the string does not include a character subset, a leaf node is configured at the end of the node of the string, an end character is configured for the leaf node, the end character is stored in the character space, and the weight corresponding to the string is stored in the weight space of the leaf node.

17. The device according to any one of claims 14 to 16, characterized in that The dictionary tree construction module is also used to obtain the first character of each business keyword and obtain the hash value corresponding to each first character; Based on the hash value of each first character, determine the subtree cluster identifier corresponding to each business keyword; distribute the business keywords and corresponding weights to the subtree cluster servers corresponding to each subtree cluster identifier, and construct a dictionary tree based on the received business keywords and corresponding weights through each subtree cluster server.

18. The device according to claim 17, characterized in that The dictionary tree construction module is also used to determine the target parameter value corresponding to each business keyword based on the hash value and a preset algorithm; determine the storage location corresponding to each business keyword based on the target parameter value; and assign a corresponding subtree cluster identifier to each business keyword according to the storage location.

19. The device according to claim 13, characterized in that The data search module is further configured to determine a search path in the target search subtree according to the direct prefix node, and traverse at least one child node corresponding to the direct prefix node according to the search path.

20. The device according to claim 19, characterized in that The data search module is also used to traverse the child nodes corresponding to the direct prefix node in the target search subtree according to the search path; obtain the weight of the current character node, and when the weight of the current character node is empty, traverse the next node of the current character node; Until the leaf node including the end character is traversed, the weight of the leaf node is obtained; The character string corresponding to the leaf node and the weight of the character string are determined according to the search path.

21. The device according to any one of claims 19 to 20, characterized in that The search suggestion word and the corresponding weight are obtained by obtaining the first character of the search character and obtaining the hash value corresponding to the first character; Determine the subtree cluster identifier corresponding to the search character based on the hash value of the first character; The search character is forwarded to a target subtree cluster server corresponding to the subtree cluster identifier, and is obtained from a pre-constructed dictionary tree based on the search character by the target subtree cluster server.

22. The device according to claim 13, characterized in that The data processing module is also used to determine the weight of the search keyword corresponding to each character string based on the weight of each character string; sort each search suggestion word in descending order based on the weight, and obtain the search suggestion word based on the sorting result and the quantity threshold; return the obtained search suggestion word to the terminal, and display the search suggestion word through the search page of the terminal according to the sorting result.

23. The device according to claim 13, characterized in that The device also includes: The dictionary tree update module is used to obtain historical business data according to a preset frequency, wherein the historical business data includes business keywords and corresponding current search popularity; determine the current weight of each business keyword according to the current search popularity of each business keyword; and update the dictionary tree based on each business keyword and the corresponding current weight.

24. The device according to claim 23, characterized in that The dictionary tree update module is also used to extract the updated business keywords and the corresponding current weights if the business keywords include updated business keywords; determine the corresponding updated characters and character sequences according to the prefix characters of the updated business keywords; and update the updated characters and the corresponding current weights in the dictionary tree based on the character sequence.

25. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.

26. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.

27. A computer program product comprising computer instructions, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Method and system for recommending user search word

    CN103150409A

  • Input prompt method and device and dictionary tree model establishing method and device

    CN103914569A