Query Request Completion Method, Device, Electronic Device, and Storage Medium
By constructing a standardized corpus and building a prefix tree, the problem of large system overhead and long request time in the query request automatic completion method is solved, and efficient user input experience is achieved with query request completion and optimization.
Patent Information
- Application Number
- CN202011476378.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-12-14
AI Technical Summary
In the prior art, the system overhead caused by the automatic completion method of query requests is large and the request time is too long, which affects the user input experience.
Construct a standardized corpus set and build a prefix tree. Query matching nodes through the prefix tree to query corpus completion query requests, reducing system overhead and request time.
It improves the reliability and efficiency of query request completion, reduces the lag during continuous input by users, and improves the user input experience.
Smart Images

Figure CN113779176B_ABST
Abstract
Description
Background Art
[0002] Query (a query request, a message sent by a search engine or database to find a specific file, website, record, or series of records in a database) auto-completion is often used in search engines. The goal is to predict the complete query during the user's input process, sort them by relevance, and recommend them to the user. By assisting the user in entering the query, it improves the user experience and avoids entering misspelled or ambiguous queries. With the rise of commercial dialogue systems, query auto-completion has also been introduced into them.
[0003] In related technologies, the main process of query auto-completion includes: after obtaining the prefix (prefix) input by the user, first recall a batch of candidate queries related to the user input from a pre-set query database through a specific recall algorithm, then perform relevance sorting on the candidate queries through a specific sorting algorithm, and finally recommend several queries with the highest rankings to the user. However, this method currently has the following defects:
[0004] Since the auto-completion process involves the retrieval of tens of millions of corpus in ElasticSearch (abbreviated as es), as well as sorting and matching the retrieval, it causes lags during the user's continuous input process, or the user has started entering the next character before the current query auto-completion result is returned, which not only affects the auto-completion effect but also the user's input experience.
[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] The purpose of the present disclosure is to provide a query request auto-completion method, device, electronic device, and computer-readable storage medium, which can at least to some extent improve the problems of large system overhead and long request time caused by the auto-completion method in related technologies.
[0007] Other features and advantages of the present disclosure will become apparent through the following detailed description, or be learned in part through the practice of the present disclosure.
[0008] According to one aspect of the present disclosure, a query request auto-completion method is provided, including: constructing a standardized corpus set of query requests; building a prefix tree based on the standardized corpus set; when obtaining the prefix of the query request input by the user, querying the node query corpus matching the prefix in the prefix tree; and auto-completing the query request based on the node query corpus.
[0009] In one embodiment, the standardized corpus for constructing the query request includes: constructing the standardized corpus based on historical query corpora and corresponding first consultation volumes, and / or standard query sentences between pre-stored users and robots and corresponding second consultation volumes.
[0010] In one embodiment, the constructing the standardized corpus based on historical query corpora and corresponding first consultation volumes, and / or standard query sentences between pre-stored users and robots and corresponding second consultation volumes includes: performing a screening operation on the historical query corpora based on preset screening conditions, and generating a first corpus based on the screening results and the first consultation volume; obtaining the corresponding standard query sentences based on preset intent classification, and generating a second corpus based on the standard query sentences and the second consultation volume;
[0011] generating the standardized corpus based on the first corpus and / or the second corpus.
[0012] In one embodiment, the performing a screening operation on the historical query corpora based on preset screening conditions, and generating a first corpus based on the screening results and the first consultation volume includes: selecting conversation logs related to the query request within a preset time period forward from the current moment; deleting stop words in the conversation logs to generate a corpus to be processed; extracting multiple types of similar query requests in the corpus to be processed based on the edit distance algorithm, and merging each type of the similar query requests to obtain multiple types of merged query requests; counting the consultation volume of each type of the similar query requests as the first consultation volume; screening the multiple types of merged query requests based on the relationship between the first consultation volume and a consultation quantity threshold, and determining the screened merged query requests as the first corpus.
[0013] In one embodiment, it further includes: generating a consultation form based on the merged query requests and the first consultation volume; the generating a second corpus based on the standard query sentences and the second consultation volume includes: performing a similarity match between the standard query sentences and the consultation form to determine the second consultation volume; generating the second corpus based on the standard query sentences and the consultation volume of the standard query sentences.
[0014] In one embodiment, the constructing a prefix tree based on the standardized corpus includes: performing word segmentation processing on the corpus in the standardized corpus to form word segmentation strings at different layers; determining the consultation times of the word segmentation strings based on the first consultation volume and / or the second consultation volume; using the word segmentation strings as edges and the consultation times of the word segmentation strings as nodes to construct the prefix tree.
[0015] In one embodiment, the step of constructing the prefix tree by using the word segmentation string as an edge and the consultation times of the word segmentation string as a node further includes: for the nodes generated by the word segmentation strings of each layer, sorting them according to the consultation times of the word segmentation strings to construct the prefix tree.
[0016] In one embodiment, constructing the prefix tree based on the standardized corpus set includes: obtaining entity information in a specified domain; extracting entity characters in the word segmentation string based on the entity information; and using the same generalization character to replace the entity characters to generate the generalized prefix tree.
[0017] In one embodiment, when obtaining the prefix of the query request input by the user, querying the node query corpus matching the prefix in the prefix tree includes: when obtaining the prefix, extracting the entity characters in the prefix based on the named entity recognition operation; using the generalization character to replace the entity characters to perform generalization processing on the entity characters; and performing a query operation in the prefix tree based on the generalization character and other characters in the prefix to obtain the corresponding node query corpus.
[0018] In one embodiment, when obtaining the prefix of the query request input by the user, querying the node query corpus matching the prefix in the prefix tree further includes: using the entity characters to replace the generalization characters in the node query corpus to complete the query request based on the replaced node query corpus.
[0019] In one embodiment, periodically obtain the updated conversation log of the query request; generate a standardized updated corpus set based on the updated conversation log and the standard query question sentence; perform generalization processing on the standardized updated corpus set to obtain a generalized corpus set; and update the prefix tree based on the generalized corpus set.
[0020] According to a second aspect of the present disclosure, there is provided a query request completion device, including: a construction module for constructing a standardized corpus set of a query request; a construction module for constructing a prefix tree based on the standardized corpus set; a query module for querying a node query corpus matching the prefix in the prefix tree when obtaining the prefix of the query request input by the user; and a completion module for completing the query request based on the node query corpus.
[0021] According to a third aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the query request completion method of any one of the above via executing the executable instructions.
[0022] According to a fourth aspect of the present disclosure, there is provided a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the query request completion method of any one of the above is implemented.
[0023] The query request completion solution provided by the embodiments of the present disclosure constructs a prefix tree according to a standardized corpus, so that when a prefix of a query request query is received, a node query corpus matching the prefix is queried in the prefix tree as the part to be completed of query, so as to complete the completion of query. The construction of the standardized corpus is beneficial to standardize the input of users, and thus can ensure the reliability of query completion. And based on the prefix retrieval of the prefix tree, it can well improve the problems of large system overhead and long request time caused by the es recall + sorting algorithm in the related art, and thus reduce the lag when the user continuously inputs, improve the completion effect, and thus improve the user input experience.
[0024] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0026] Figure 1 A schematic diagram showing the structure of a query request completion system in an embodiment of the present disclosure;
[0027] Figure 2 A flowchart showing a query request completion method in an embodiment of the present disclosure;
[0028] Figure 3 A flowchart showing another query request completion method in an embodiment of the present disclosure;
[0029] Figure 4 A flowchart showing still another query request completion method in an embodiment of the present disclosure;
[0030] Figure 5 A structural diagram showing a prefix tree in an embodiment of the present disclosure;
[0031] Figure 6 A flowchart showing yet another query request completion method in an embodiment of the present disclosure;
[0032] Figure 7Flowchart showing another query request completion method according to an embodiment of the present disclosure;
[0033] Figure 8 Modular flowchart showing a query request completion method according to an embodiment of the present disclosure;
[0034] Figure 9 Schematic diagram showing a query request completion device according to an embodiment of the present disclosure;
[0035] Figure 10 Schematic diagram showing an electronic device according to an embodiment of the present disclosure. Detailed implementation manners
[0036] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.
[0037] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0038] The solution provided in this application constructs a prefix tree according to a standardized corpus, so that when the prefix of a query request query is received, the node query corpus matching the prefix is queried in the prefix tree as the part to be completed of the query, so as to complete the completion of the query. The construction of the standardized corpus is beneficial to standardize the input of users, and thus can ensure the reliability of query completion. Moreover, based on the prefix retrieval of the prefix tree, it can well improve the problems of large system overhead and long request time caused by the es recall + sorting algorithm in the related art, thereby reducing the lag that occurs when users continuously input, improving the completion effect, and thus enhancing the user's input experience.
[0039] For ease of understanding, several terms related to this application will be explained first below.
[0040] Query: That is, a query. Specifically, it is a message sent by a search engine or a database in order to find a specific file, website, record, or a series of records in a database. In this disclosure, Query refers to a query request.
[0041] Prefix: That is, a prefix. Specifically, it is the intermediate text generated by a user during the process of inputting a query, and is usually the first half of the final query.
[0042] ElasticSearch (es) is specifically a search and data analysis engine.
[0043] Named Entity Recognition (NER) is an important basic tool in application fields such as information extraction, question answering systems, syntactic analysis, and machine translation, and occupies an important position in the process of the practical application of natural language processing technology. Generally speaking, the task of named entity recognition is to identify three major categories (entity category, time category, and number category) and seven minor categories (person name, organization name, place name, time, date, currency, and percentage) of named entities in the text to be processed.
[0044] The prefix tree, also known as the word search tree, Trie tree, is a tree structure and a variant of the hash tree. Its typical application is for statistics, sorting, and storing a large number of strings (but not limited to strings), so it is often used by search engine systems for text word frequency statistics. Its advantages are: using the common prefix of strings to reduce the query time, minimizing unnecessary string comparisons, and having a higher query efficiency than the hash tree.
[0045] Edit distance algorithm, edit distance (Minimum Edit Distance, MED), also known as the Levenshtein distance, refers to the minimum number of edit operations required to convert one string into another between two strings. The allowed edit operations include: replacing one character with another character (substitution, s), inserting a character (insert, i), or deleting a character (delete, d).
[0046] The solution provided by the embodiments of this application involves technologies such as face recognition and machine learning, and is specifically described through the following embodiments.
[0047] Figure 1 The structural schematic diagram of a query request completion system in the embodiments of this disclosure is shown, including a plurality of terminals 120 and a server cluster 140.
[0048] The terminal 120 can be a mobile terminal such as a mobile phone, a game console, a tablet computer, an e-book reader, smart glasses, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a smart home device, an AR (Augmented Reality) device, a VR (Virtual Reality) device, etc. Alternatively, the terminal 120 can also be a personal computer (PC), such as a laptop computer and a desktop computer, etc.
[0049] Among them, an application program for providing query request completion can be installed in the terminal 120.
[0050] The terminal 120 is connected to the server cluster 140 through a communication network. Optionally, the communication network is a wired network or a wireless network.
[0051] The server cluster 140 is a single server, or consists of several servers, or is a virtualization platform, or is a cloud computing service center. The server cluster 140 is used to provide background services for the query request completion application program. Optionally, the server cluster 140 undertakes the main computing work, and the terminal 120 undertakes the secondary computing work; or, the server cluster 140 undertakes the secondary computing work, and the terminal 120 undertakes the main computing work; or, the terminal 120 and the server cluster 140 adopt a distributed computing architecture for collaborative computing.
[0052] In some alternative embodiments, the server cluster 140 is used to store the query request completion model, etc.
[0053] Optionally, the clients of the application programs installed in different terminals 120 are the same, or the clients of the application programs installed on two terminals 120 are clients of the same type of application program on different control system platforms. Based on the differences in the terminal platforms, the specific forms of the clients of the application program can also be different. For example, the client of the application program can be a mobile phone client, a PC client, or a World Wide Web (Web) client, etc.
[0054] Those skilled in the art can know that the number of the above terminals 120 can be more or less. For example, there can be only one of the above terminals, or there can be dozens or hundreds of the above terminals, or even more. The embodiments of the present application do not limit the number and device types of the terminals.
[0055] Optionally, the system can also include a management device ( Figure 1(not shown), the management device is connected to the server cluster 140 through a communication network. Optionally, the communication network is a wired network or a wireless network.
[0056] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is usually the Internet, but can also be any network, including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network). In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent the data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some of the links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.
[0057] Next, each step in the query request completion method in the present exemplary embodiment will be described in more detail with reference to the accompanying drawings and embodiments.
[0058] Figure 2 The flowchart of a query request completion method in an embodiment of the present disclosure is shown. The method provided by the embodiment of the present disclosure can be executed by any electronic device with computing and processing capabilities, such as Figure 1 the terminal 120 and / or the server cluster 140 in. In the following illustrative examples, the terminal 120 is used as the execution subject for illustrative purposes.
[0059] As Figure 2 shown, the terminal 120 executes the query request completion method, including the following steps:
[0060] Step S202, construct a standardized corpus for the query request.
[0061] Among them, the standardized corpus refers to a corpus that is easy for the robot executing the query request query to understand. The construction of the standardized corpus is beneficial to standardize the user's input and reduce the difficulty of subsequent robot intention recognition and response.
[0062] Step S204, construct a prefix tree based on the standardized corpus.
[0063] Among them, the prefix tree constructed based on the standardized corpus can optimize the retrieval path, thereby reducing the time consumption of query completion.
[0064] Step S206, when the prefix of the query request input by the user is obtained, query the node query corpus matching the prefix in the prefix tree.
[0065] Among them, based on the result of retrieving the prefix tree, the query instruction is completed.
[0066] Step S208, complete the query request based on the node query corpus.
[0067] Among them, completing the query request based on the node query corpus is specifically to display it in the input box of the prefix in a specified manner.
[0068] In this embodiment, by constructing a prefix tree according to the standardized corpus, when the prefix of the query request query is received, the node query corpus matching the prefix is queried in the prefix tree as the part to be completed of the query, so as to complete the query completion. The construction of the standardized corpus is beneficial to standardize the user's input, and thus can ensure the reliability of query completion.
[0069] Furthermore, based on the prefix retrieval of the prefix tree, it can well improve the problems of large system overhead and long request time caused by the es recall + sorting algorithm in the related technology, thereby reducing the lag when the user inputs continuously, improving the completion effect, and thus enhancing the user's input experience.
[0070] In one embodiment, step S202 of constructing the standardized corpus of the query request includes: constructing the standardized corpus of the query request based on the historical query corpus and the corresponding first consultation volume, and / or the standard query questions between the pre-stored user and the robot and the corresponding second consultation volume.
[0071] Among them, the historical query corpus represents the relevant corpus input by the collected users, and the first consultation volume represents the number of query frequencies of each historical query corpus.
[0072] The standard query questions between users and the robot refer to the predefined intention classifications of users in the intelligent customer service system and the corresponding standard question forms. These standard question forms are sentences that have been manually screened, conform to grammar norms, and are easy to understand. If users directly ask these questions, the robot can smoothly complete the intention recognition and response processes. The second consultation volume represents the frequency of each standard query question.
[0073] In this embodiment, constructing a standardized corpus set based on historical query corpora and standard query questions can recommend standard queries to users, standardize user inputs, and reduce the difficulty of subsequent robot intention recognition and response. At the same time, since these are high-frequency and standardized questions, this data set naturally filters out long-tail, non-standard expressions, swear words, and sentences that are particularly short or long.
[0074] For example, a standard query refers to a sentence in a chatbot where the standard question can be correctly recognized for its intention. For a chatbot constructed according to a text classification system, one classification is "price protection", and the corresponding standard question is "I want to apply for price protection".
[0075] In one embodiment, constructing a standardized corpus set for query requests based on historical query corpora and the corresponding first consultation volume, and / or the standard query questions between users and the robot and the corresponding second consultation volume includes:
[0076] Performing a screening operation on the historical query corpus based on preset screening conditions, and generating a first corpus set based on the screening results and the first consultation volume.
[0077] Among them, the preset screening conditions are conditions related to time and consultation volume, such as the historical query corpus being the conversation logs of the recent week or month, the conversation logs with a consultation volume greater than a certain threshold, or the conversation logs ranked among the top few in terms of consultation volume, etc.
[0078] Obtaining the corresponding standard query questions based on the preset intention classifications, and generating a second corpus set based on the standard query questions and the second consultation volume; generating a standardized corpus set based on the first corpus set and / or the second corpus set.
[0079] Among them, the historical query corpus includes statements similar to the standard query questions, and the consultation volume of the standard query questions is statistically calculated based on the consultation volume of the similar statements.
[0080] Such as Figure 3As shown, in one embodiment, a screening operation is performed on historical query corpora based on preset screening conditions, a first corpus set is generated based on the screening results and the first consultation volume, and corresponding standard query sentences are obtained based on preset intent classification, so as to generate a second corpus set based on the standard query sentences and the second consultation volume; A specific implementation manner of generating a standardized corpus set based on the first corpus set and / or the second corpus set includes:
[0081] Step S302, select the conversation logs related to the query request within a preset time duration forward from the current moment.
[0082] Step S304, delete the stop words in the conversation logs to generate a corpus set to be processed.
[0083] Step S306, extract multiple types of similar query requests in the corpus set to be processed based on the edit distance algorithm, and merge each type of similar query requests to obtain multiple types of merged query requests.
[0084] Step S308, count the consultation volume of each type of similar query request as the first consultation volume.
[0085] Step S310, screen the multiple types of merged query requests based on the relationship between the first consultation volume and the consultation quantity threshold, and determine the screened merged query requests as the first corpus set.
[0086] Step S312, generate a consultation form based on the merged query requests and the first consultation volume.
[0087] Among them, the historical query corpus is specifically the historical high-frequency corpus. Normally, for high-frequency questions, the robot can understand and answer them well. Therefore, when preparing data, high-frequency corpus is preferred first. Specifically, the following method can be used to find high-frequency corpus:
[0088] Select the conversation system logs of the recent week, and all queries are organized by day dimension.
[0089] Perform stop word removal processing on all queries, such as: ah|ne|this|that|these|those|here|there.
[0090] Set the threshold of the edit distance algorithm to be greater than 0.85 as equal, find similar queries for merging, and calculate the consultation volume of the queries. This step can obtain the query_consultation_map, that is, the consultation form, where the key is the processed similar query and the value is the first consultation volume of the query.
[0091] In the query_consultation_map, select the data with the top 50% of the consultation volume as alternatives, add all the data together, remove the duplicate content, and use it as the screening result to generate the first corpus set.
[0092] Step S314: Match the standard query sentence with the consultation form for similarity to determine the second consultation volume.
[0093] Step S316: Generate a second corpus based on the consultation volume of the standard query sentence and the standard query sentence.
[0094] Step S318: Perform a deduplication operation on the basis of the first corpus and the second corpus to generate a standardized corpus.
[0095] Specifically, the intelligent customer service system generally predefines the intention classification of users and gives each classification a standard query sentence. These standard query sentences are manually screened, conform to grammatical norms, and are easy-to-understand sentences. If the user directly asks these questions, the robot can smoothly complete the intention recognition and response process.
[0096] Directly find these standard query sentences in the robot database. For example: "The refund has not arrived", "Is there a charge for opening a white bar?", "Usage restrictions of coupons", which can provide about 2,000 standard query sentences.
[0097] For each standard query sentence, a consultation volume data is given, which will be used in the subsequent construction of the prefix tree. Here, we will use each standard query sentence in turn to perform a similarity match with each element in the query_consultation_map obtained in the first step. Similarly, if the edit distance is greater than 0.85, it is considered that the two are equal. In this way, it is equivalent to using the real consultation situation of the online data to construct a consultation volume information for the offline standard query sentence to obtain the second corpus.
[0098] Add the first corpus and the second corpus together, and after deduplication, obtain a standardized data set.
[0099] In this embodiment, since the standardized corpus is generated based on high-frequency and standardized corpora, this data set has already filtered out unqualified sentences such as long tails, unstandardized expressions, swear words, extremely short or extremely long sentences. Therefore, by constructing a standardized corpus, to recommend a standard query to the user, through a completion operation, the user inputs a standard query, so as to standardize the content input by the user, and further reduce the difficulty of subsequent robot intention recognition and response.
[0100] As Figure 4 shown, step S204: Constructing a prefix tree based on the standardized corpus includes:
[0101] Step S402: Perform word segmentation processing on the corpora in the standardized corpus to form word segmentation strings of different layers.
[0102] Step S404: Determine the consultation times of the word segmentation strings based on the first consultation volume and / or the second consultation volume.
[0103] Step S406: Use the word segmentation strings as edges and the consultation times of the word segmentation strings as nodes to construct a prefix tree.
[0104] Step S408: For the nodes generated by the word segmentation strings of each layer, sort them according to the consultation times of the word segmentation strings to generate a prefix tree.
[0105] Specifically, all the corpora in the standardized corpus are segmented, and the consultation times of each word segmentation string are counted. For example: "When will my things be delivered", in the query_consultation_map in 3.2.2, the consultation times cot_1 of this query can be queried; "The things I bought yesterday have already dropped in price today", query the consultation times cot_2 of this query; "My mobile phone screen is broken", query the consultation times cot_3 of this query. As shown in Table 1, the consultation times of the word segmentation string "I" are cot_1 + cot_2 + cot_3; the consultation times of the word segmentation string "my" are cot_1 + cot_3; the consultation times of the word segmentation string "I just" are cot_2.
[0106] Table 1
[0107]
[0108] Take all the word segmentation strings as the edges of the prefix tree, record the consultation times of the word segmentation strings in each child node, and arrange all the child nodes of each layer from left to right in descending order of consultation times. The more to the left, the more times it appears, which is convenient for searching. The root nodes of all are the times that each standard corpus sentence in the corpus appears. The prefix tree constructed based on the above process is as Figure 5 shown.
[0109] In this embodiment, by constructing a prefix tree, the retrieval of the prefix each time is realized. Due to the advantages of the prefix tree itself in terms of search efficiency, it can well avoid the problems of large system overhead and long request time brought by the traditional two-step method of es recall + sorting algorithm.
[0110] As Figure 6 shown, in one embodiment, constructing a prefix tree based on a standardized corpus includes:
[0111] Step S602: Obtain entity information in a specified domain.
[0112] Step S604: Extract entity characters in the word segmentation string based on the entity information.
[0113] Step S606: Use the same generalization character to replace the entity characters to generate a generalized prefix tree.
[0114] Implement the generalization process of the corpus through NER technology to further narrow the overall corpus and also identify more user questions.
[0115] Next, take the e-commerce field as the specified field and specifically describe the generalization process.
[0116] The NER technology in the e-commerce field can identify the products in the query. For example, "computer", "mobile phone", etc. can be identified as product entities and marked as [PRODSORT], which is the generalization character.
[0117] For the user question "Why hasn't the mobile phone I bought arrived yet", it is generalized to "Why hasn't the [PRODSORT] I bought arrived yet"; the user question "I need to return my computer" is generalized to "I need to return my [PRODSORT]" and saved in the dataset.
[0118] Corresponding to the online process, an NER module will be added to identify the product name. When the user actually asks "the mobile phone I bought", it will first be generalized to "the [PRODSORT] I bought" and then retrieved;
[0119] For the retrieved result, corresponding replacement will be made and then shown to the user. Here, [PRODSORT] will be replaced with "product", and finally "Why hasn't the product I bought arrived yet" will be returned.
[0120] In this embodiment, through the generalization process of the corpus, the probability of the retrieval result being empty due to entity inconsistency can be effectively reduced, and through flexible matching operations, it is ensured to provide the user with the desired information.
[0121] As Figure 7 shown, in one embodiment, step S206: When obtaining the prefix of the query request input by the user, query the nodes in the prefix tree that match the prefix. The query corpus includes:
[0122] Step S702: When obtaining the prefix, extract the entity characters in the prefix based on the named entity recognition operation.
[0123] Step S704: Use the generalization character to replace the entity characters to perform generalization processing on the entity characters.
[0124] Step S706: Based on the generalization character and other characters in the prefix, perform a query operation in the prefix tree to obtain the corresponding node query corpus.
[0125] Step S708: Use the entity characters to replace the generalization characters in the node query corpus to complete the query request based on the replaced node query corpus.
[0126] Specifically, when the user enters the line to consult a question and inputs a prefix "prefix", the query auto-completion service is triggered, which is divided into the following steps:
[0127] The prefix first performs entity recognition to obtain entity characters. For example, when the user inputs "my mobile phone", it will be generalized to "my [PRODSORT]", which is the generalized character.
[0128] Based on the generalization result, directly retrieve in the prefix tree to find the child node with the largest consultation volume. The edge of this child node is the returned query information.
[0129] To avoid all recommended queries being similar sentences, we have some techniques in the specific retrieval. For example Figure 6 As shown, when the user inputs "you", if directly looking for the top 3 child nodes with the largest consultation volume, they are likely to be all under the node "you all". To make the returned information more diverse, here we will respectively select the child node with the largest consultation volume under each of the three nodes "you all", "hello", and "you are".
[0130] The retrieved data is directly returned to the user. If there is an entity recognition result, it is replaced first and then returned.
[0131] In this embodiment, the query auto-completion based on the prefix tree is different from the traditional two-step process of recall + ranking. The whole process only needs to query the prefix tree once to return the result, which greatly simplifies the previous online processing process.
[0132] Due to the natural advantage of the tree model in querying, the overall query time can be controlled at the millisecond level. In actual experiments, the single query time for data at the hundred-thousand level can be controlled at about 2ms.
[0133] In one embodiment, regularly obtain the updated conversation log of the query request; generate a standardized updated corpus based on the updated conversation log and the standard query sentences; perform generalization processing on the standardized updated corpus to obtain a generalized corpus; update the prefix tree based on the generalized corpus.
[0134] Due to the traditional es recall + ranking strategy, limited by the huge es corpus and the long and manually intervened corpus preparation process, it is almost impossible to achieve automatic update of the corpus.
[0135] The present disclosure proposes a query auto-completion method based on the prefix tree. The whole process from corpus standardization, corpus generalization to the construction of the prefix tree is more lightweight, and the prefix tree can be updated automatically. The specific automatic update process is as follows:
[0136] Regularly obtain the dialogue system logs. Generally scheduled at midnight every day to automatically obtain all user queries for the current day.
[0137] Data standardization processing is generally performed using data from the past 10 days. In actual situations, data from a longer time period can also be used for calculation. First, calculate the query_consultation_map. The consultation volume is calculated as cot_all = a*cot_today + b*cot_yesterday + … + i*cot_last_9day + j*last_10day, where a to j are set as a = 1.0, b = 0.9, c = 0.8, d = 0.7, e = 0.6, f = 0.5, g = 0.4, h = 0.3, i = 0.2, j = 0.1. The closer the data is to the current time, the greater the weight it occupies in the consultation volume. The calculated consultation volume data can not only well reflect the current popular questions of user consultations but also retain the consultation information of the past week to a certain extent.
[0138] The data generalization process refers to the above generalization process. For the current data processing of less than 100,000, the overall time consumption is about 1 hour.
[0139] The prefix tree construction process refers to the above construction process. The time to construct the prefix tree is very fast. For data at the 100,000 level, it can be completed in a few minutes.
[0140] In this embodiment, through automatic data acquisition and prefix tree construction, the overall process can be completed within 2 hours. It is completely possible to utilize the time when there are no user incoming calls in the early morning to automatically update and optimize the prefix tree, keeping the underlying data up-to-date at all times. Furthermore, it can ensure that with the changes in the chatbot business, there are still sufficient recommendable queries, thereby ensuring the click-through rate of user recommended queries.
[0141] As Figure 8 shown, the entire query request completion process can be divided into an online part and an offline part.
[0142] Among them, the online part includes step S802, obtaining the prefix input by the user; step S804, prefix entity recognition; step S806, prefix tree search; step S808, returning the query.
[0143] The offline part includes step S810, constructing a standardized corpus set; step S812, corpus generalization; step S814, prefix tree construction; step S816, regularly obtaining logs. In step S816, steps S810 to S814 are repeatedly executed to achieve the update of the prefix tree.
[0144] It should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for restrictive purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0145] Those skilled in the art can understand that various aspects of the present invention can be implemented as a system, a method, or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuits", "modules", or "systems" here.
[0146] Next, refer to Figure 9 to describe the query request completion device 900 according to this embodiment of the present invention. Figure 9 The shown query request completion device 900 is only an example and should not bring any limitations to the functions and usage scopes of the embodiments of the present invention.
[0147] The query request completion device 900 is presented in the form of a hardware module. The components of the query request completion device 900 may include, but are not limited to: a construction module 902 for constructing a standardized corpus set of query requests; a building module 904 for building a prefix tree based on the standardized corpus set; a query module 906 for querying node query corpora matching the prefix in the prefix tree when obtaining the prefix of the query request input by the user; and a completion module 908 for completing the query request based on the node query corpora.
[0148] In one embodiment, the construction module 902 is further configured to: construct the standardized corpus set based on historical query corpora and corresponding first consultation volumes, and / or standard query questions between pre-stored users and robots and corresponding second consultation volumes.
[0149] In one embodiment, the construction module 902 is further configured to: perform a screening operation on historical query corpora based on preset screening conditions, generate a first corpus set based on the screening results and the first consultation volume; obtain the corresponding standard query questions based on preset intention classifications, so as to generate a second corpus set based on the standard query questions and the second consultation volume; and generate the standardized corpus set based on the first corpus set and / or the second corpus set.
[0150] In one embodiment, the construction module 902 is further configured to: select conversation logs related to the query request within a preset duration forward from the current moment; delete stop words in the conversation logs to generate a corpus to be processed; extract multiple types of similar query requests in the corpus to be processed based on the edit distance algorithm, and merge each type of the similar query requests to obtain multiple types of merged query requests; count the consultation volume of each type of the similar query requests as the first consultation volume; screen the multiple types of merged query requests based on the relationship between the first consultation volume and the consultation quantity threshold, and determine the screened merged query requests as the first corpus.
[0151] In one embodiment, the construction module 902 is further configured to: generate a consultation form based on the merged query request and the first consultation volume; the generating the second corpus based on the standard query sentence and the second consultation volume includes: performing a similarity match between the standard query sentence and the consultation form to determine the second consultation volume; generating the second corpus based on the standard query sentence and the consultation volume of the standard query sentence.
[0152] In one embodiment, the construction module 904 is further configured to: perform word segmentation processing on the corpus in the standardized corpus to form word segmentation strings at different layers; determine the consultation times of the word segmentation strings based on the first consultation volume and / or the second consultation volume; use the word segmentation strings as edges and the consultation times of the word segmentation strings as nodes to construct the prefix tree.
[0153] In one embodiment, the construction module 904 is further configured to: sort the nodes generated by the word segmentation strings at each layer according to the consultation times of the word segmentation strings to construct the prefix tree.
[0154] In one embodiment, the construction module 904 is further configured to: obtain entity information in a specified domain; extract entity characters in the word segmentation strings based on the entity information; use the same generalization character to replace the entity characters to generate a generalized prefix tree.
[0155] In one embodiment, the query module 906 is further configured to: when obtaining the prefix, extract the entity characters in the prefix based on the named entity recognition operation; use the generalization character to replace the entity characters to perform generalization processing on the entity characters; perform a query operation in the prefix tree based on the generalization character and other characters in the prefix to obtain the corresponding node query corpus.
[0156] In one embodiment, the query module 906 is further configured to: use the entity characters to replace the generalization characters in the node query corpus to complete the query request based on the replaced node query corpus.
[0157] In one embodiment, it further includes an update module 910, and the update module 910 is configured to: periodically obtain the updated conversation log of the query request; generate a standardized updated corpus set based on the updated conversation log and the standard query sentence; perform generalization processing on the standardized updated corpus set to obtain a generalized corpus set; update the prefix tree based on the generalized corpus set.
[0158] The following will refer to Figure 10 to describe the electronic device 1000 according to this embodiment of the present invention. Figure 10 The shown electronic device 1000 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0159] As Figure 10 shown, the electronic device 1000 is presented in the form of a general-purpose computing device. The components of the electronic device 1000 may include but are not limited to: the above-mentioned at least one processing unit 1010, the above-mentioned at least one storage unit 1020, and a bus 1030 connecting different system components (including the storage unit 1020 and the processing unit 1010).
[0160] Among them, the storage unit stores program codes, and the program codes can be executed by the processing unit 1010, so that the processing unit 1010 executes the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification. For example, the processing unit 1010 can execute steps S202, S204 to S208 as shown in Figure 2 and other steps defined in the query request completion method of the present disclosure.
[0161] The storage unit 1020 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 10201 and / or a cache storage unit 10202, and may further include a read-only storage unit (ROM) 10203.
[0162] The storage unit 1020 may further include a program / utilities 10204 having a set (at least one) of program modules 10205. Such program modules 10205 include but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0163] The bus 1030 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.
[0164] The electronic device 1000 may also communicate with one or more external devices 1060 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and may also communicate with one or more devices that enable a user to interact with the electronic device, and / or communicate with any device (such as a router, a modem, etc.) that enables the electronic device 1000 to communicate with one or more other computing devices. Such communication may be carried out through the input / output (I / O) interface 1050. Moreover, the electronic device 1000 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 1050. As shown in the figure, the network adapter 1050 communicates with other modules of the electronic device 1000 through the bus 1030. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in combination with the electronic device, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0165] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0166] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above method of this specification is stored. In some possible implementation manners, various aspects of the present invention may also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification.
[0167] The program product for implementing the above method according to the embodiments of the present invention may adopt a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.
[0168] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0169] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0170] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).
[0171] It should be noted that although several modules or units of the devices for action execution are mentioned in the foregoing detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described modules or units may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by multiple modules or units.
[0172] In addition, although the steps of the methods in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0173] Those skilled in the art can easily understand from the description of the above embodiments that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0174] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
Claims
1. A query request completion method, characterized in that, Including: Constructing a standardized corpus for query requests, including: constructing the standardized corpus based on historical query corpora and corresponding first consultation volumes, and / or standard query questions between pre-stored users and robots and corresponding second consultation volumes, wherein the first consultation volume represents the query frequency of each historical query corpus, and the second consultation volume represents the query frequency of each standard query question; Constructing a prefix tree based on the standardized corpus, including: performing word segmentation on the corpora in the standardized corpus to form word segmentation strings at different levels; determining the consultation times of the word segmentation strings based on the first consultation volume and / or the second consultation volume; using the word segmentation strings as edges and the consultation times of the word segmentation strings as nodes to construct the prefix tree; When obtaining the prefix of the query request input by the user, querying the node query corpus matching the prefix in the prefix tree; Completing the query request based on the node query corpus.
2. The query request completion method according to claim 1, wherein The constructing the standardized corpus based on historical query corpora and corresponding first consultation volumes, and / or standard query questions between pre-stored users and robots and corresponding second consultation volumes includes: Performing a screening operation on the historical query corpora based on preset screening conditions, and generating a first corpus based on the screening results and the first consultation volume; Obtaining the corresponding standard query questions based on preset intention classification, and generating a second corpus based on the standard query questions and the second consultation volume; Generating the standardized corpus based on the first corpus and / or the second corpus.
3. The query request completion method according to claim 2, wherein The performing a screening operation on the historical query corpora based on preset screening conditions, and generating a first corpus based on the screening results and the first consultation volume includes: Selecting the conversation logs related to the query request within a preset time period forward from the current moment; Deleting the stop words in the conversation logs to generate a corpus to be processed; Extracting multiple types of similar query requests in the corpus to be processed based on the edit distance algorithm, and merging each type of the similar query requests to obtain multiple types of merged query requests; Counting the consultation volume of each type of the similar query requests as the first consultation volume; Screening the multiple types of merged query requests based on the relationship between the first consultation volume and the consultation quantity threshold, and determining the screened merged query requests as the first corpus.
4. The query request completion method according to claim 3, characterized in that Also including: Generating a consultation form based on the merged query requests and the first consultation volume; The generating a second corpus based on the standard query questions and the second consultation volume includes: Performing similarity matching between the standard query questions and the consultation form to determine the second consultation volume; Generating the second corpus based on the standard query questions and the consultation volume of the standard query questions.
5. The query request completion method according to claim 1, wherein The using the word segmentation strings as edges and the consultation times of the word segmentation strings as nodes to construct the prefix tree further includes: Sorting the nodes generated by the word segmentation strings at each level according to the consultation times of the word segmentation strings to construct the prefix tree.
6. The query request completion method according to claim 1, wherein, The constructing a prefix tree based on the standardized corpus includes: Obtain entity information in a specified domain; Extract entity characters in the word segmentation string based on the entity information; Use the same generalization character to replace the entity characters to generate the generalized prefix tree.
7. The query request completion method according to claim 6, characterized in that When obtaining the prefix of the query request input by the user, querying for nodes in the prefix tree that match the prefix. The query corpus includes: When obtaining the prefix, extract the entity characters in the prefix based on named entity recognition operations; Use the generalization character to replace the entity characters to perform generalization processing on the entity characters; Based on the generalization character and other characters in the prefix, perform a query operation in the prefix tree to obtain the corresponding node query corpus.
8. The query request completion method according to claim 7, characterized in that When obtaining the prefix of the query request input by the user, querying for nodes in the prefix tree that match the prefix. The query corpus also includes: Use the entity characters to replace the generalization characters in the node query corpus to complete the query request based on the replaced node query corpus.
9. The query request completion method according to any one of claims 1 to 4, characterized in that Further includes: Regularly obtain the updated conversation log of the query request; Generate a standardized updated corpus set based on the updated conversation log and the standard query question sentences; Perform generalization processing on the standardized updated corpus set to obtain a generalized corpus set; Update the prefix tree based on the generalized corpus set.
10. A query request completion device, characterized in that, Includes: A construction module for constructing a standardized corpus set of query requests, including: constructing the standardized corpus set based on historical query corpora and corresponding first consultation volumes, and / or standard query question sentences and corresponding second consultation volumes pre-stored between users and robots, where the first consultation volume represents the query frequency of each historical query corpus, and the second consultation volume represents the query frequency of each standard query question sentence; A construction module for constructing a prefix tree based on the standardized corpus set, including: performing word segmentation processing on the corpora in the standardized corpus set to form word segmentation strings at different layers; determining the consultation times of the word segmentation strings based on the first consultation volume and / or the second consultation volume; using the word segmentation strings as edges and the consultation times of the word segmentation strings as nodes to construct the prefix tree; A query module for querying for node query corpora in the prefix tree that match the prefix when obtaining the prefix of the query request input by the user; A completion module for completing the query request based on the node query corpus.
11. An electronic device, characterized in that, Includes: A processor; And A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the query request completion method according to any one of claims 1 to 9 by executing the executable instructions.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the query request completion method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Universal special word recognition method and system based on mode expansion
CN111159990A
Entity linking method and device
CN111401049A