Question and answer guiding method, device and system
By performing word segmentation processing on the query statements entered by the user, gradually determining the candidate word collection, the problem that the recommendation information in the prior art is not consistent with the actual needs of the user, and a more accurate recommendation effect is achieved.
Patent Information
- Application Number
- CN202410029874.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-08
- Publication Date
- 2025-07-11
AI Technical Summary
The existing search engine recommendation methods only provide recommendation information based on the last word or word entered by the user, resulting in a large deviation from the query information that the user actually needs to input, making it difficult to meet the actual needs of the user.
By performing word segmentation processing on the query statement input by the user, the candidate word set of each word segment is obtained, and the candidate word set is gradually determined until the recommended word is obtained. The candidate word set of each word segmentation in the query statement is considered, thereby reducing the deviation between the recommended word and the query information that the user actually needs to input.
By gradually selecting the candidate word set, the deviation between the recommended word and the query information that the user actually needs to input is reduced, and the actual needs of the user are met.
Smart Images

Figure CN120296114A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of search technologies, and in particular, to a question-and-answer guiding method, apparatus, and system. Background Art
[0002] In the related art, when a user queries information through a search engine, the user only needs to input some words or some texts, and the search engine can give recommended complete words according to the partial information input by the user. The user selects the information to be input from the recommended information, so as to obtain the complete query information. However, the existing recommendation method gives relevant recommended information only according to the last word or the last character input by the user. In this way, the recommended information given is quite different from the query information that the user actually needs to input, and it is difficult to meet the actual needs of the user. Summary of the Invention
[0003] To overcome the problems existing in the related art, the present disclosure provides a question-and-answer guiding method, apparatus, and system.
[0004] According to a first aspect of an embodiment of the present disclosure, there is provided a question-and-answer guiding method, including:
[0005] Performing word segmentation processing on a query statement input by a user to obtain N word segments; where N is an integer greater than 1;
[0006] Obtaining a candidate word set including the first word segment among the N word segments;
[0007] For the current word segment among the N word segments, obtaining a candidate word set including the current word segment from the candidate word set including the first word segment until obtaining a first candidate word set including the (N - 1)-th word segment; the current word segment is any other word segment among the N word segments except the first word segment and the N-th word segment, and the first word segment is the previous word segment of the current word segment among the N word segments;
[0008] Based on the first candidate word set, determining a second candidate word set corresponding to the N-th word segment;
[0009] Selecting a recommended word from the second candidate word set and outputting the recommended word.
[0010] In some embodiments of the present disclosure, the determining, based on the first candidate word set, a second candidate word set corresponding to the N-th word segment includes:
[0011] In the case of determining that at least one of the first candidate word sets includes the N-th word segment, obtaining a second candidate word set including the N-th word segment from the first candidate word set;
[0012] In the case that it is determined that the Nth word segment is not included in the first candidate word set, obtain the word prefixes of each candidate word in the first candidate word set, and determine a second candidate word set corresponding to the Nth word segment based on the word prefixes.
[0013] In some embodiments of the present disclosure, obtaining a second candidate word set including the Nth word segment from the first candidate word set includes:
[0014] In the case that it is determined that at least one of the first candidate word sets includes the Nth word segment, determine the type of the Nth word segment;
[0015] Select at least one candidate word different from the type from the second candidate word set to obtain the recommended word.
[0016] In some embodiments of the present disclosure, each candidate word set includes multiple candidate words and the type of each candidate word; before performing word segmentation processing on the query statement input by the user to obtain a word segmentation sequence including N word segments, the method further includes:
[0017] For each candidate word among the multiple candidate words, summarize all types of each candidate word to obtain type information corresponding to the candidate word;
[0018] Among them, determining the type of the Nth word segment includes:
[0019] Determine the type information corresponding to the candidate word identical to the Nth word segment to obtain the type of the Nth word segment.
[0020] In some embodiments of the present disclosure, determining a second candidate word set corresponding to the Nth word segment based on the word prefixes includes:
[0021] Select a target word prefix identical to the Nth word segment from the word prefixes;
[0022] Determine the first candidate word corresponding to the target word prefix to obtain the second candidate word set.
[0023] In some embodiments of the present disclosure, before performing word segmentation processing on the query statement input by the user, the method further includes:
[0024] For each candidate word in all candidate word sets for question-and-answer guidance, summarize the candidate word sets including the candidate word to obtain set information corresponding to the candidate word;
[0025] Among them, obtaining a candidate word set including the first word segment among the N word segments includes:
[0026] Determine the set information corresponding to the second candidate word that is the same as the first participle, and obtain a plurality of candidate word sets including the first participle according to the set information corresponding to the second candidate word;
[0027] Among them, obtaining the candidate word set including the current participle from the candidate word set including the first participle includes:
[0028] Determine the set information corresponding to the third candidate word that is the same as the current participle, and obtain the candidate word set including the current participle from the candidate word set including the first participle according to the set information corresponding to the third candidate word.
[0029] In some embodiments of the present disclosure, before performing participle processing on the query statement input by the user, the method further includes:
[0030] Obtain all candidate words in the entire candidate word set for question and answer guidance;
[0031] Add the all candidate words to the participle dictionary for performing participle processing on the query statement input by the user.
[0032] According to a second aspect of the embodiments of the present disclosure, there is provided a question and answer guidance device, including:
[0033] A participle unit, configured to perform participle processing on a query statement input by a user to obtain N participles; where N is an integer greater than 1;
[0034] A first acquisition unit, configured to acquire a candidate word set including the first participle among the N participles;
[0035] A second acquisition unit, configured to, for the current participle among the N participles, acquire a candidate word set including the current participle from the candidate word set including the first participle until a first candidate word set including the (N - 1)-th participle is obtained; the current participle is any other participle except the first participle and the N-th participle among the N participles, and the first participle is the previous participle of the current participle among the N participles;
[0036] A determination unit, configured to determine a second candidate word set corresponding to the N-th participle based on the first candidate word set;
[0037] An output unit, configured to select a recommended word from the second candidate word set and output the recommended word.
[0038] In some embodiments of the present disclosure, the determination unit is specifically configured to:
[0039] When it is determined that the Nth word segment is included in at least one of the first candidate word sets, obtain a second candidate word set including the Nth word segment from the first candidate word set;
[0040] When it is determined that the Nth word segment is not included in the first candidate word set, obtain the word prefixes of each candidate word in the first candidate word set, and determine the second candidate word set corresponding to the Nth word segment based on the word prefixes.
[0041] In some embodiments of the present disclosure, the output unit is specifically configured to:
[0042] When it is determined that the Nth word segment is included in at least one of the first candidate word sets, determine the type of the Nth word segment;
[0043] Select at least one candidate word different from the type from the second candidate word set to obtain the recommended word.
[0044] In some embodiments of the present disclosure, each candidate word set includes multiple candidate words and the type of each candidate word, and the apparatus further includes:
[0045] A first summarizing unit, configured to summarize all types of each candidate word among the multiple candidate words to obtain type information corresponding to the candidate word;
[0046] Wherein, the determining unit is specifically configured to:
[0047] Determine the type information corresponding to the candidate word that is the same as the Nth word segment to obtain the type of the Nth word segment.
[0048] In some embodiments of the present disclosure, the determining unit is specifically configured to:
[0049] Select a target word prefix that is the same as the Nth word segment from the word prefixes;
[0050] Determine the first candidate word corresponding to the target word prefix to obtain the second candidate word set.
[0051] In some embodiments of the present disclosure, the apparatus further includes:
[0052] A second summarizing unit, configured to summarize the candidate word sets including the candidate word for each candidate word in all candidate word sets for question-and-answer guidance to obtain set information corresponding to the candidate word;
[0053] Wherein, the first obtaining unit is specifically configured to:
[0054] Determine the set information corresponding to the second candidate word that is the same as the first participle, and obtain a plurality of candidate word sets including the first participle according to the set information corresponding to the second candidate word;
[0055] Wherein, the second obtaining unit is specifically configured to:
[0056] Determine the set information corresponding to the third candidate word that is the same as the current participle, and obtain a candidate word set including the current participle from the candidate word set including the first participle according to the set information corresponding to the third candidate word.
[0057] In some embodiments of the present disclosure, the apparatus further includes:
[0058] A third obtaining unit, configured to obtain all candidate words in all candidate word sets for question and answer guidance;
[0059] An adding unit, configured to add the all candidate words to a participle dictionary for performing participle processing on the query statement input by the user.
[0060] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including:
[0061] A memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the method described in any one of the first aspect is implemented.
[0062] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any one of the first aspect is implemented.
[0063] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any one of the first aspect is implemented.
[0064] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: performing word segmentation on the query statement input by the user to obtain N word segments, obtaining a candidate word set including the first word segment among the N word segments, for the current word segment among the N word segments, obtaining a candidate word set including the current word segment from the candidate word set including the first word segment, until obtaining a first candidate word set including the (N-1)th word segment, based on the first candidate word set, determining a second candidate word set corresponding to the Nth word segment, selecting a recommended word from the second candidate word set, and outputting the recommended word. That is, by gradually selecting to obtain a candidate word set including the first (N-1) word segments, and then determining a second candidate word set corresponding to the Nth word segment according to the candidate word set, a candidate word set considering each word segment in the query statement can be obtained, thereby narrowing the range of the second candidate word set, and further reducing the deviation between the output recommended word and the query information that the user actually needs to input, meeting the actual needs of the user.
[0065] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0067] Figure 1 is a flowchart of a question-and-answer guiding method shown according to an exemplary embodiment.
[0068] Figure 2 is a schematic structural diagram of a data graph shown according to an embodiment of the present disclosure.
[0069] Figure 3 is a flowchart of another question-and-answer guiding method shown according to an exemplary embodiment.
[0070] Figure 4 is a flowchart of yet another question-and-answer guiding method shown according to an exemplary embodiment.
[0071] Figure 5 is a flowchart of a question-and-answer guiding method shown according to another exemplary embodiment.
[0072] Figure 6 is a flowchart of a question-and-answer guiding method shown according to still another exemplary embodiment.
[0073] Figure 7 is a block diagram of a question-and-answer guiding device shown according to an exemplary embodiment.
[0074] Figure 8It is a block diagram of a device for a question-and-answer guiding method shown according to an exemplary embodiment. Detailed implementation manners
[0075] Here, the exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0076] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments, and are not intended to limit the embodiments of the present disclosure. The singular forms "a" and "the" used in the embodiments of the present disclosure and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0077] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the embodiments of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "when" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0078] In addition, various forms of processes shown in the embodiments of the present disclosure can be used, and steps can be reordered, added, or deleted. For example, the steps described in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved, and no limitations are imposed herein.
[0079] It should be noted that in the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information and other processes all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0080] In the related art, when a user queries information through a search engine, the user only needs to input some words or some texts, and the search engine can give recommended complete words based on the partial information input by the user. The user selects the information to be input from the recommended information, so as to obtain the complete query information. However, the existing recommendation method gives relevant recommended information only based on the last word or the last character input by the user. The recommended information given in this way has a large deviation from the query information that the user actually needs to input, and it is difficult to meet the actual needs of the user.
[0081] To solve the above problems, the present disclosure provides a question-and-answer guiding method, apparatus, and system.
[0082] Figure 1 is a flowchart of a question-and-answer guiding method shown according to an exemplary embodiment. As Figure 1 shown, it should be noted that the question-and-answer guiding method of the embodiments of the present disclosure is applied to a question-and-answer guiding apparatus. As Figure 1 shown, the method may include the following steps:
[0083] Step 101, perform word segmentation processing on the query statement input by the user to obtain N word segments.
[0084] Wherein, in the embodiments of the present disclosure, N is an integer greater than 1.
[0085] In some embodiments of the present disclosure, the above query statement may be text data directly input by the user, or may also be text data obtained by recognizing the voice or picture input by the user.
[0086] In an embodiment of the present disclosure, after receiving the query statement input by the user, perform word segmentation processing on the query statement to obtain N word segments.
[0087] For example, performing word segmentation processing on the query statement input by the user may obtain a vocabulary list word_list as [word_1, word_2, …, word_N].
[0088] It should be noted that the above query statement may be a complete statement or an incomplete statement.
[0089] In one embodiment, a word segmentation method may be used to perform word segmentation processing on the query statement to obtain N word segments.
[0090] In another embodiment, multiple word segmentation methods may be used to perform word segmentation processing on the query statement respectively to obtain the word segmentation results corresponding to each word segmentation method, and the word segmentation result that more meets the actual requirements is selected as the above N word segments according to a preset rule.
[0091] Step 102, obtain a candidate word set including the first word segment among the N word segments.
[0092] It should be noted that the candidate word set may include multiple candidate words for guiding the user to ask questions and get answers.
[0093] In some embodiments of the present disclosure, each candidate word set may be a data graph, such as Figure 2As shown, the data graph can be multiple data tables including different types of nodes. The data tables include the following nodes: data table ID, metric name, dimension name, dimension value, and synonym. Taking data table 1 with data table ID as table_1 as an example, the metric name can be cost or output; the dimension name can be model name or product channel; the dimension value can be the value corresponding to the dimension name. For example, the dimension values corresponding to the model name include model 1 and model 2, and the dimension values corresponding to the product channel include channel 1 and channel 2; the synonym can be a synonym of the metric name. For example, the synonyms corresponding to cost include synonym 1 and synonym 2, or can be synonyms of the dimension value. For example, the dimension values corresponding to model 1 include synonym 3 and synonym 4. In addition, the data table also includes the association relationships between different nodes, such as: <data table ID, metric name>, <data table ID, dimension name>, <dimension name, dimension value>, <dimension value, synonym>, <dimension name, synonym>, <metric name, synonym>.
[0094] Among them, a metric is a data stream in which the metric value changes over time. For example, the output changes over time, and the outputs corresponding to different time nodes are different. The information corresponding to the metric can be used to generate a query result for a query statement. For example, if the query statement is "the output change trend of xx mobile phones in the past week", then a line chart of the output change in the past week can be generated based on the output data; a dimension is data that does not change. For example, "model 1" corresponding to the model name does not change.
[0095] Before step 101, the above data graph can be established in advance using a database. The database can be a Neo4j knowledge base.
[0096] It can be understood that the above obtained N word segments have a sequential order, that is, the order of the N word segments in the query statement.
[0097] In the embodiments of the present disclosure, in accordance with the above order, the first word segment among the N word segments can be determined, and a candidate word segment set including the above first word segment can be screened from multiple pre-stored candidate word segment sets.
[0098] Step 103, for the current word segment among the N word segments, obtain a candidate word segment set including the current word segment from the candidate word segment set including the first word segment until a first candidate word segment set including the (N - 1)-th word segment is obtained.
[0099] Among them, in the embodiments of the present disclosure, the current word segment is any other word segment among the N word segments except the first word segment and the N-th word segment, and the first word segment is the previous word segment of the current word segment among the N word segments.
[0100] In the embodiments of the present disclosure, the following processing is sequentially performed in the order of the N word segments in the query statement: obtaining a candidate word set including the current word segment from multiple candidate sets including the first word segment, that is, screening out a candidate word set including the current word segment from multiple candidate word sets corresponding to the previous word segment, so as to gradually narrow the range of candidate words until a first candidate word set including the (N - 1)-th word segment is obtained.
[0101] Step 104: Based on the first candidate word set, determine a second candidate word set corresponding to the N-th word segment.
[0102] In the embodiments of the present disclosure, a second candidate word set for providing recommended words for the N-th word segment can be determined according to the first candidate word set.
[0103] Step 105: Select a recommended word from the second candidate word set and output the recommended word.
[0104] It should be noted that there can be multiple such recommended words.
[0105] In an embodiment of the present disclosure, a recommended word can be randomly selected from the second candidate word set and output.
[0106] In another embodiment of the present disclosure, a recommended word that meets a preset condition can also be selected from the second candidate word set according to a pre-rule.
[0107] According to the question-and-answer guiding method of the embodiments of the present disclosure, by performing word segmentation processing on a query statement input by a user to obtain N word segments, obtaining a candidate word set including the first word segment among the N word segments, for the current word segment among the N word segments, obtaining a candidate word set including the current word segment from the candidate word set including the first word segment until a first candidate word set including the (N - 1)-th word segment is obtained, based on the first candidate word set, determining a second candidate word set corresponding to the N-th word segment, selecting a recommended word from the second candidate word set and outputting the recommended word, that is, by gradually selecting to obtain a candidate word set including the first N - 1 word segments, and then determining a second candidate word set corresponding to the N-th word segment according to this candidate word set, so as to obtain a candidate word set that takes into account each word segment in the query statement, thereby narrowing the range of the second candidate word set, and further reducing the deviation between the output recommended word and the query information that the user actually needs to input, meeting the actual needs of the user.
[0108] Next, the implementation process of determining a second candidate word set corresponding to the N-th word segment based on the first candidate word set will be introduced in detail.
[0109] Figure 3 It is a flowchart of another question-and-answer guiding method shown according to an exemplary embodiment. As Figure 3 shown, based on the above embodiments,Figure 1 The implementation process of determining the second candidate word set corresponding to the Nth word segment based on the first candidate word set in step 104 may include the following steps:
[0110] Step 301: when it is determined that at least one first candidate word set includes the Nth word segment, a second candidate word set including the Nth word segment is obtained from the first candidate word set.
[0111] It should be noted that since the query sentence input by the user may be a complete sentence or an incomplete sentence, the last segmentation obtained by segmenting the query sentence may be a complete word or an incomplete word. If the last segmentation is a complete word, the segmentation can be found in the first candidate word set. If the last segmentation is not a complete word, the segmentation cannot be directly found in the first candidate word set.
[0112] In the embodiment of the present disclosure, when it is determined that at least one first candidate word set includes the Nth word segment, it is indicated that the Nth word segment is a complete word, and a second candidate word set including the Nth word segment is directly selected from the first candidate word set.
[0113] Step 302: When it is determined that the first candidate word set does not include the Nth word segment, obtain the word prefix of each candidate word in the first candidate word set, and determine the second candidate word set corresponding to the Nth word segment based on the word prefix.
[0114] In the embodiment of the present disclosure, when it is determined that at least one first candidate word set includes the Nth word segment, it is indicated that the Nth word segment is an incomplete word segment and may be a part of a candidate word or several candidate words in the first candidate word set. Therefore, the word prefix of each candidate word in the first candidate word set can be obtained, and the second candidate word set corresponding to the Nth word segment can be determined based on the word prefix.
[0115] For example, the Nth participle is "第". If this participle is not found in the first candidate word set, it means that "第" may be the word prefix of some candidate words in the second candidate word set, such as "第一周" and "第一星期". The second candidate word set corresponding to the Nth participle can be determined by searching the word prefix of each candidate word in the first candidate word set.
[0116] In some embodiments of the present disclosure, the implementation process of step 302 of determining the second candidate word set corresponding to the Nth word participle based on the word prefix may include the following steps: selecting a target word prefix that is the same as the Nth word participle from the word prefixes, determining the first candidate word corresponding to the target word prefix, and obtaining the second candidate word set.
[0117] In an embodiment of the present disclosure, a target word prefix that is the same as the Nth word segment is searched for from the word prefixes of multiple candidate words in the first candidate word set, and the first candidate word corresponding to the target word prefix is determined. All the first candidate words corresponding to the target word prefixes are used as the second candidate word set.
[0118] It should be noted that steps 301 and 302 do not distinguish the order of execution.
[0119] According to the Q&A guiding method of the embodiments of the present disclosure, when it is determined that at least one first candidate word set includes the Nth word segment, a second candidate word set including the Nth word segment is obtained from the first candidate word set. When it is determined that the first candidate word set does not include the Nth word segment, the word prefixes of each candidate word in the first candidate word set are obtained, and a second candidate word set corresponding to the Nth word segment is determined based on the word prefixes. Thus, on the basis of narrowing the range of candidate words, different methods for determining the second candidate word set are matched by identifying the integrity of the last word segment.
[0120] Next, the implementation process of obtaining the second candidate word set including the Nth word segment from the first candidate word set will be introduced in detail.
[0121] Figure 4 is a flowchart of another Q&A guiding method shown according to an exemplary embodiment. As Figure 4 shown, based on the above embodiments, Figure 1 The implementation process of selecting the recommended word from the second candidate word set in step 105 in
[0122] Step 401, when it is determined that at least one first candidate word set includes the Nth word segment, determine the type of the Nth word segment.
[0123] As an example of a possible implementation manner, the type of each candidate word can be pre-stored in the candidate word set. The above types can be index names, dimension names, dimension values, synonyms of dimension values, or synonyms of index names.
[0124] In the embodiments of the present disclosure, based on the data graph proposed in some embodiments of the present disclosure, the type of the Nth word segment is determined.
[0125] In some embodiments of the present disclosure, before Figure 1 performing word segmentation processing on the query statement input by the user in step 101 in
[0126] to obtain a word segmentation sequence including N word segments, the method further includes the following steps:
[0127] Among them, in the embodiments of the present disclosure, each candidate word set includes multiple candidate words and the type of each candidate word.
[0128] It can be understood that since the same candidate word can appear in different candidate word sets, the type of the same candidate word in each candidate word set may be different. Therefore, there can be multiple types for each subsequent word segmentation. If the type of the word segmentation is determined each time by looking up the types of the candidate words identical to the word segmentation in all the subsequent candidate word sets, the efficiency is low.
[0129] Therefore, for each candidate word among multiple candidate words, the types of all candidate words can be summarized in advance to obtain the type information corresponding to the candidate word, thereby improving the search efficiency.
[0130] For example, the types of word_i include index name, dimension value, etc. The types of word_i can be expressed as word_i_type = [type_i1, type_i2,..., type_iM], where M represents the number of types.
[0131] Among them, the implementation process of determining the type of the Nth word segmentation in step 301 may include the following steps:
[0132] Determine the type information corresponding to the candidate word identical to the Nth word segmentation to obtain the type of the Nth word segmentation.
[0133] In the embodiments of the present disclosure, determine the type information corresponding to the candidate word identical to the Nth word segmentation, and determine all types of the Nth word segmentation according to the above type information.
[0134] Step 402: Select at least one candidate word with a different type from the second candidate word set to obtain a recommended word.
[0135] It can be understood that usually in the same query statement, the types of two adjacent words are different. Therefore, according to the type of the Nth word segmentation, the candidate words with different types from the Nth word segmentation can be determined as the candidate word range of the (N + 1)th word segmentation.
[0136] In the embodiments of the present disclosure, select at least one candidate word with a different type from the second candidate word set to obtain a recommended word for providing question - answering guidance to the user.
[0137] According to the question - answering guidance method of the embodiments of the present disclosure, when it is determined that the Nth word segmentation is included in at least one first candidate word set, determine the type of the Nth word segmentation, select at least one candidate word with a different type from the second candidate word set to obtain a recommended word, so as to ensure that the recommended word provided to the user is a word of a different type from the last word segmentation, and improve the rationality of the question - answering guidance.
[0138] Figure 5 is a flowchart of a Q&A guiding method shown according to another exemplary embodiment. As Figure 5 shown, the method includes the following steps:
[0139] Step 501, for each candidate word in the entire set of candidate words for Q&A guiding, summarize the set of candidate words including the candidate word to obtain the set information corresponding to the candidate word.
[0140] It should be noted that the same candidate word can appear in multiple different sets of candidate words. If it is necessary to determine which sets of candidate words the first token belongs to every time obtaining the set of candidate words including the current token from the set of candidate words including the first token, the efficiency is relatively low.
[0141] Therefore, for each candidate word in the entire set of candidate words for Q&A guiding, the sets of candidate words including the candidate word can be pre-summarized to obtain the set information corresponding to the candidate word. When it is necessary to determine the set of candidate words corresponding to the token, it can be directly found through the set information, effectively improving the search efficiency.
[0142] Step 502, perform tokenization on the query statement input by the user to obtain N tokens.
[0143] Wherein, N is an integer greater than 1.
[0144] It should be noted that in the embodiments of the present disclosure, the above step 502 can be implemented in any one of the embodiments of the present disclosure. The embodiments of the present disclosure do not make any limitations on this and will not be elaborated any further.
[0145] Step 503, determine the set information corresponding to the second candidate word that is the same as the first token, and obtain multiple sets of candidate words including the first token according to the set information corresponding to the second candidate word.
[0146] In the embodiments of the present disclosure, the set information may be the identification information of the set of candidate words. For example, when the set of candidate words is a data table in the data graph proposed in the embodiments of the present disclosure, the set information may be the table name of the data table or the ID of the data table.
[0147] In the embodiments of the present disclosure, the first token is used as the second candidate word, the set information corresponding to the second candidate word is searched, the set information of all sets of candidate words including the first token is determined according to the set information corresponding to the second candidate word, and multiple sets of candidate words including the first token are obtained according to the set information.
[0148] Step 504: For the current token among the N tokens, determine the set information corresponding to the third candidate token identical to the current token. According to the set information corresponding to the third candidate token, obtain the candidate token set including the current token from the candidate token set including the first token until the first candidate token set including the (N - 1)-th token is obtained.
[0149] Wherein, the current token is any token other than the first token and the N-th token among the N tokens, and the first token is the token preceding the current token among the N tokens.
[0150] In the embodiments of the present disclosure, for the current token among the N tokens, the current token can be used as the third candidate token, and the set information corresponding to the third candidate token can be searched to obtain the set information of the candidate token set including the current token. Based on the above set information, the candidate token set including the current token is obtained from the candidate token set including the first token until the first candidate token set including the (N - 1)-th token is obtained.
[0151] For example, it can be searched in the data graph which data table IDs each word appears in. The data table can be one or more. For example, the i-th token word_i exists in data tables such as table_201 and table_301 at the same time. Here, the set information corresponding to word_i can be expressed as word_i_table = [table_i1, table_i2, …, table_iK], where K is the number of tables.
[0152] As a possible example, when the sentence input by the user is a question-and-answer query, the vocabulary list word_list including j tokens obtained after word segmentation is [word_1, word_2, …, word_j]. Here, the query can be a complete sentence or an incomplete sentence.
[0153] According to the association relationship information dict_relation, search in table_ok to obtain the data table including word_1, and obtain the data table ID to which word_1 belongs, that is, word_1_table. The formula f for this process is:
[0154] word_1_table = f(word_1, table_ok, relation_dict)
[0155] Wherein, table_ok is the data table constraint range of word_1. Since word_1 is the first token, table_ok here is the full amount of data tables, that is, all the data tables used for question-and-answer guidance.
[0156] According to the information in dict_relation, search for the data table including word_2 in word_1_table to obtain the data table ID to which word_2 belongs, and obtain word_2_table. This process is as follows:
[0157] word_2_table = f(word2, word_1_table, relation_dict)
[0158] And so on until the last word word_j is obtained. It should be noted here that word_j may be a complete word or an incomplete word. If it is determined that word_j is not a complete word, use word-initial guidance (i.e., step 301 of the embodiments of the present disclosure). If it is determined that word_j is a complete word, use adjacent-word guidance (i.e., step 302 of the embodiments of the present disclosure). The formula d for this process is:
[0159] is_full_word = s(word_j, dict_relation),
[0160] where the judgment result is_full_word is 1 or 0. 1 indicates that word_j is a complete word, and 0 indicates that word_j is not a complete word.
[0161] In some embodiments of the present disclosure, the association relationship information dict_relation can store the relationships between each candidate word and the data table ID (i.e., the set information proposed in some embodiments of the present disclosure), node type, dimension name, and lexical prefix. For example, dict_relation stores the set information word_i_table, node type information word_i_type, dimension name information word_i_dimension, and lexical prefix information word_i_prefix of the i-th candidate word word_i. Among them, word_i_table = [table_i1, table_i2,..., table_iK], where K is the number of tables; word_i_type = [type_i1, type_i2,..., type_iM], where M represents the number of types; word_i_dimension = [dimension_i1, dimension_i2,..., dimension_iP], where P is the number of dimension names, and word_i_prefix = [prefix_i1, prefix_i2,..., prefix_iQ], where Q represents the number of prefixes.
[0162] Step 505: Based on the first candidate word set, determine the second candidate word set corresponding to the Nth word segment.
[0163] As an example, if word_j is not a complete word, then the word-initial guiding method is adopted. The formula f for this process is:
[0164] word_j_prefix = g(word_j, table_{j - 1})
[0165] where word_j_prefix is the second candidate word set, and table_{j - 1} is the first candidate word set, that is, the candidate word set corresponding to the (j - 1)th word segment.
[0166] As another example, if word_j is a complete word, adjacent word guiding is used. According to the information in dict_relation, the node type information of word_j is obtained as word_j_type. The formula g for this process is:
[0167] word_j_type = l(word_j, dict_relation)
[0168] Randomly select a candidate word with a type different from that in word_j_type in word_j_table as the guiding range word_i_neighbor of the adjacent word of word_j, that is, the guiding range of the next word of word_j. The formula h for this process is:
[0169] word_i_neighbor = h(word_j, word_j_table, word_j_type)
[0170] Step 506: Select a recommended word from the second candidate word set and output the recommended word.
[0171] It should be noted that in the embodiments of the present disclosure, the above Step 506 can be implemented in any one of the embodiments of the present disclosure. The embodiments of the present disclosure do not make any limitations in this regard and will not be elaborated further.
[0172] According to the question-and-answer guiding method of the embodiments of the present disclosure, for each candidate word in the entire candidate word set for question-and-answer guiding, the candidate word sets including the candidate word are summarized to obtain the set information corresponding to the candidate word, so that the candidate word set corresponding to the current word segment can be quickly determined, effectively improving the efficiency of determining the second candidate word set, and further improving the efficiency of providing recommended words for users.
[0173] Figure 6is a flowchart of a question and answer guiding method shown according to another exemplary embodiment. As Figure 6 shown, based on the above embodiment, before the step 101 of Figure 1 performing word segmentation on the query statement input by the user, the method further includes the following steps:
[0174] Step 601, obtaining all candidate words in the entire candidate word set for question and answer guiding.
[0175] Step 602, adding all candidate words to the word segmentation dictionary for performing word segmentation on the query statement input by the user.
[0176] It should be noted that in order to be able to segment each word when performing word segmentation on the query statement input by the user, all candidate words in the entire candidate word set for question and answer guiding can be pre-loaded into the word segmentation dictionary for performing word segmentation on the query statement input by the user by using a word segmentation tool, so as to improve the accuracy of word segmentation processing.
[0177] In some embodiments of the present disclosure, the above word segmentation tool may be the Jieba word segmentation tool.
[0178] According to the question and answer guiding method of the embodiments of the present disclosure, all candidate words in the entire candidate word set for question and answer guiding are obtained, and all candidate words are added to the word segmentation dictionary for performing word segmentation on the query statement input by the user. It ensures that all candidate words in the entire candidate word set for question and answer guiding can be segmented, and improves the accuracy of word segmentation processing.
[0179] Figure 7 is a block diagram of a question and answer guiding device shown according to an exemplary embodiment. Referring to Figure 7 , the device includes a word segmentation unit 701, a first obtaining unit 702, a second obtaining unit 703, a determining unit 704, and an output unit 705.
[0180] Among them, the word segmentation unit 701 is used to perform word segmentation on the query statement input by the user to obtain N word segments; where N is an integer greater than 1;
[0181] The first obtaining unit 702 is used to obtain a candidate word set including the first word segment among the N word segments;
[0182] The second obtaining unit 703 is used to, for the current word segment among the N word segments, obtain a candidate word set including the current word segment from the candidate word set including the first word segment until a first candidate word set including the (N-1)th word segment is obtained; the current word segment is any other word segment among the N word segments except the first word segment and the Nth word segment, and the first word segment is the previous word segment of the current word segment among the N word segments;
[0183] A determination unit 704, configured to determine a second candidate word set corresponding to the Nth word segment based on a first candidate word set;
[0184] An output unit 705, configured to select a recommended word from the second candidate word set and output the recommended word.
[0185] In some embodiments of the present disclosure, the determination unit 704 is specifically configured to:
[0186] When it is determined that at least one first candidate word set includes the Nth word segment, obtain a second candidate word set including the Nth word segment from the first candidate word set;
[0187] When it is determined that the first candidate word set does not include the Nth word segment, obtain the word prefixes of each candidate word in the first candidate word set, and determine a second candidate word set corresponding to the Nth word segment based on the word prefixes.
[0188] In some embodiments of the present disclosure, the output unit 705 is specifically configured to:
[0189] When it is determined that at least one first candidate word set includes the Nth word segment, determine the type of the Nth word segment;
[0190] Select at least one candidate word different from the type from the second candidate word set to obtain a recommended word.
[0191] In some embodiments of the present disclosure, each candidate word set includes multiple candidate words and the type of each candidate word. The question-and-answer guiding device further includes:
[0192] A first summarization unit, configured to summarize all types of each candidate word among multiple candidate words to obtain type information corresponding to the candidate word;
[0193] Wherein, the determination unit 704 is specifically configured to:
[0194] Determine the type information corresponding to the candidate word identical to the Nth word segment to obtain the type of the Nth word segment.
[0195] In some embodiments of the present disclosure, the determination unit 704 is specifically configured to:
[0196] Select a target word prefix identical to the Nth word segment from the word prefixes;
[0197] Determine the first candidate word corresponding to the target word prefix to obtain a second candidate word set.
[0198] In some embodiments of the present disclosure, the question-and-answer guiding device further includes:
[0199] A second summarization unit, configured to summarize, for each candidate word in the entire candidate word set for question-and-answer guidance, the candidate word set including the candidate word, to obtain the set information corresponding to the candidate word;
[0200] Among them, the first obtaining unit 702 is specifically configured to:
[0201] Determine the set information corresponding to the second candidate word that is the same as the first segmented word, and according to the set information corresponding to the second candidate word, obtain multiple candidate word sets including the first segmented word;
[0202] Among them, the second obtaining unit 703 is specifically configured to:
[0203] Determine the set information corresponding to the third candidate word that is the same as the current segmented word, and according to the set information corresponding to the third candidate word, obtain the candidate word set including the current segmented word from the candidate word set including the first segmented word.
[0204] In some embodiments of the present disclosure, the question-and-answer guidance device further includes:
[0205] A third obtaining unit, configured to obtain all candidate words in the entire candidate word set for question-and-answer guidance;
[0206] An adding unit, configured to add all candidate words to the word segmentation dictionary for performing word segmentation processing on the query statement input by the user.
[0207] According to the question-and-answer guidance device of the embodiments of the present disclosure, word segmentation processing is performed on the query statement input by the user to obtain N segmented words, and a candidate word set including the first segmented word among the N segmented words is obtained. For the current segmented word among the N segmented words, a candidate word set including the current segmented word is obtained from the candidate word set including the first segmented word, until a first candidate word set including the (N - 1)-th segmented word is obtained. Based on the first candidate word set, a second candidate word set corresponding to the N-th segmented word is determined, and a recommended word is selected from the second candidate word set and output. That is, by gradually selecting to obtain the candidate word set including the first (N - 1) segmented words, and then determining the second candidate word set corresponding to the N-th segmented word according to the candidate word set, a candidate word set considering each segmented word in the query statement is obtained, thereby narrowing the range of the second candidate word set, and further reducing the deviation between the output recommended word and the query information actually required to be input by the user, meeting the actual needs of the user.
[0208] Figure 8 It is a block diagram of a device 800 for a question-and-answer guidance method shown according to an exemplary embodiment. For example, the device 800 may be an electronic device, such as a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0209] Reference Figure 8 Referring to Figure 8 , device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0210] The processing component 802 generally controls the overall operation of the device 800, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-described methods. In addition, the processing component 802 may include one or more modules to facilitate interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate interaction between the multimedia component 808 and the processing component 802.
[0211]
[0212] The memory 804 is configured to store various types of data to support the operation of the device 800. Examples of such data include instructions for any application or method operating on the device 800, contact data, phone book data, messages, pictures, videos, and the like. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a disk, or an optical disc.
[0213] The power component 806 provides power to the various components of the device 800. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 800.The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operation mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0214] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the device 800 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0215] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0216] The sensor component 814 includes one or more sensors for providing a status assessment of various aspects of the device 800. For example, the sensor component 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and the keypad of the device 800. The sensor component 814 can also detect a change in the position of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and the temperature change of the device 800. The sensor component 814 can include a proximity sensor that is configured to detect the presence of nearby objects without any physical contact. The sensor component 814 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0217] The communication component 816 is configured to facilitate wired or wireless communication between the device 800 and other devices. The device 800 may access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0218] In an exemplary embodiment, the device 800 may be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described method.
[0219] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as the memory 804 including instructions, is also provided. The above instructions may be executed by the processor 820 of the device 800 to complete the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0220] In an exemplary embodiment, a computer program product including a computer program is also provided. The computer program implements the above-described method when executed by the processor 820 of the device 800.
[0221] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include known common knowledge or conventional technical means in the technical field not disclosed herein. The specification and embodiments are only to be considered exemplary, and the true scope and spirit of the invention are pointed out by the following claims.
[0222] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. A question-and-answer guiding method, characterized in that, Comprising: Performing word segmentation on the query statement input by the user to obtain N word segments; where N is an integer greater than 1; Obtaining a candidate word set including the first word segment among the N word segments; For the current word segment among the N word segments, obtaining a candidate word set including the current word segment from the candidate word set including the first word segment until obtaining a first candidate word set including the (N - 1)-th word segment; the current word segment is any other word segment among the N word segments except the first word segment and the N-th word segment, and the first word segment is the previous word segment of the current word segment among the N word segments; Based on the first candidate word set, determining a second candidate word set corresponding to the N-th word segment; Selecting a recommended word from the second candidate word set and outputting the recommended word.
2. The Q&A guiding method according to claim 1, characterized in that, The determining the second candidate word set corresponding to the N-th word segment based on the first candidate word set includes: In the case of determining that at least one of the first candidate word sets includes the N-th word segment, obtaining a second candidate word set including the N-th word segment from the first candidate word set; In the case of determining that the first candidate word set does not include the N-th word segment, obtaining the word prefixes of each candidate word in the first candidate word set, and determining a second candidate word set corresponding to the N-th word segment based on the word prefixes.
3. The Q&A guiding method according to claim 2, wherein The selecting a recommended word from the second candidate word set includes: In the case of determining that at least one of the first candidate word sets includes the N-th word segment, determining the type of the N-th word segment; Selecting at least one candidate word different from the type from the second candidate word set to obtain the recommended word.
4. The Q&A guiding method according to claim 3, wherein Each candidate word set includes multiple candidate words and the type of each candidate word; Before performing word segmentation on the query statement input by the user to obtain a word segmentation sequence including N word segments, the method further includes: For each candidate word among the multiple candidate words, summarizing all types of each candidate word to obtain type information corresponding to the candidate word; Wherein, the determining the type of the N-th word segment includes: Determining the type information corresponding to the candidate word identical to the N-th word segment to obtain the type of the N-th word segment.
5. The Q&A guiding method according to claim 2, wherein The determining the second candidate word set corresponding to the N-th word segment based on the word prefixes includes: Selecting a target word prefix identical to the N-th word segment from the word prefixes; Determining the first candidate word corresponding to the target word prefix to obtain the second candidate word set.
6. The Q&A guiding method according to claim 1, wherein Before performing word segmentation on the query statement input by the user, the method further includes: For each candidate word in the entire candidate word set for question and answer guidance, summarizing the candidate word sets including the candidate word to obtain set information corresponding to the candidate word; Wherein, the obtaining the candidate word set including the first word segment among the N word segments includes: Determining the set information corresponding to the second candidate word identical to the first word segment, and according to the set information corresponding to the second candidate word, obtaining multiple candidate word sets including the first word segment; Among them, obtaining the candidate word set including the current participle from the candidate word set including the first participle includes: Determining the set information corresponding to the third candidate word identical to the current participle, and obtaining the candidate word set including the current participle from the candidate word set including the first participle according to the set information corresponding to the third candidate word.
7. The Q&A guiding method according to claim 1, wherein Before performing participle processing on the query statement input by the user, the method further includes: Obtaining all candidate words in the entire candidate word set for question-and-answer guidance; Adding the all candidate words to the participle dictionary for performing participle processing on the query statement input by the user.
8. A question-and-answer guiding device, characterized in that, Including: A participle unit for performing participle processing on the query statement input by the user to obtain N participles; where N is an integer greater than 1; A first obtaining unit for obtaining the candidate word set including the first participle among the N participles; A second obtaining unit for, for the current participle among the N participles, obtaining the candidate word set including the current participle from the candidate word set including the first participle until obtaining the first candidate word set including the (N - 1)-th participle; the current participle is any other participle except the first participle and the N-th participle among the N participles, and the first participle is the previous participle of the current participle among the N participles; A determining unit for determining the second candidate word set corresponding to the N-th participle based on the first candidate word set; An output unit for selecting a recommended word from the second candidate word set and outputting the recommended word.
9. An electronic device, characterized in that, Including: A memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the method described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1 to 7 is implemented.