Recommended method, device, computer equipment and storage medium for extended questions
By generating a template library and matching the target extension template, the candidate extension question is constructed to determine the recommended extension question, which solves the problem of manual annotation dependency in the prior art and achieves more efficient extension questions.
Patent Information
- Application Number
- CN202110848835.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-27
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-07-27
AI Technical Summary
The prior art relies on a large number of manual annotations when recommending extension questions, resulting in high labor costs and low recommendation accuracy.
By generating a template library based on the historical dialogue corpus, obtaining the target extension template matching the standard question to be extended, and constructing candidate extension questions, the recommended extension questions are finally determined.
Reduces the dependence of manual annotation, reduces labor costs, and improves the accuracy of extended questions recommendations.
Smart Images

Figure CN113688636B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of natural language processing, and in particular to a method, apparatus, computer device and storage medium for recommending extended questions. Background Art
[0002] With the rapid development of the financial industry, commercial banks can provide customers with a variety of standardized financial products and services (such as deposits, housing loans, consumer loans, etc.). When a large number of customers use these financial products, they often have a lot of questions. As a result, the customer service system receives a large number of customer calls every day.
[0003] At present, the intelligent customer service system converts the customer's voice into text (Audio Speech Recognition, ASR), and then uses NLP (Natural Language Processing) technology to classify the customer's intentions. Then, the customer service system provides users with different services and feedback for different intentions. In order to improve the accuracy of customer intention recognition, a large number of samples (i.e. standard questions) need to be written in advance for each intention. However, the number of manually written standard questions is limited, so it is necessary to recommend a large number of relevant samples (i.e. extended questions) based on the standard questions.
[0004] Related technologies usually use the regular expression (or template) method, which involves manually sorting out several common high-frequency user expression pattern templates, and then replacing the templates accordingly based on the key terms in the current text, thereby achieving extended question recommendations. However, this method relies on a large amount of manual annotation, which results in high labor costs and low recommendation accuracy. Summary of the invention
[0005] Based on this, it is necessary to provide a recommended method, device, computer equipment and storage medium for the above technical problems, which can reduce the cost of manual expansion.
[0006] A method for extending a question, the method comprising:
[0007] Generate a template library according to a historical dialogue corpus, wherein the historical dialogue corpus includes a plurality of historical dialogue corpora, and the template library includes at least one extended question template corresponding to the historical dialogue corpus;
[0008] Acquire a target extension question template matching the standard question to be extended from the template library;
[0009] Constructing a candidate extension question according to the standard question to be extended and the target extension question template;
[0010] A recommended extension question corresponding to the standard question to be extended is determined from the candidate extension questions.
[0011] In one embodiment, generating a template library based on a historical dialogue corpus includes:
[0012] Performing character type prediction processing on historical dialogue corpus in the historical dialogue corpus to obtain a type label sequence corresponding to the historical dialogue corpus, wherein the type label sequence includes the character type corresponding to each character in the historical dialogue corpus;
[0013] According to the character types corresponding to the characters, the historical dialogue corpus is segmented to obtain the words and sentences to be replaced in the historical dialogue corpus and the semantic types corresponding to the words and sentences to be replaced;
[0014] Using a placeholder corresponding to the semantic type of the to-be-replaced phrase to replace the to-be-replaced phrase in the historical dialogue corpus to obtain an extended question template;
[0015] A template library is constructed based on the extended question model corresponding to the historical dialogue corpus.
[0016] In one embodiment, the segmenting of the historical dialogue corpus according to the character type corresponding to each character to obtain the words and sentences to be replaced in the historical dialogue corpus and the semantic types corresponding to the words and sentences to be replaced includes:
[0017] Traversing the historical dialogue corpus, and when it is determined that the character type corresponding to the currently traversed character corresponds to the first semantic type, determining the currently traversed character as the first character, and continuing to traverse the next character;
[0018] If a second character whose character type corresponds to the second semantic type or the empty type is traversed, the first character is divided into a phrase to be replaced, and the phrase to be replaced corresponds to the first semantic type;
[0019] The first semantic type is any semantic type among the semantic types, and the second semantic type is any semantic type among the semantic types except the first semantic type.
[0020] In one embodiment, the step of acquiring a target extension question template matching the standard question to be extended from the template library includes:
[0021] Performing character type prediction processing on the standard question to be expanded to obtain a type label sequence corresponding to the standard question to be expanded, wherein the type label sequence includes the character type corresponding to each character in the standard question to be expanded;
[0022] According to the character type corresponding to each of the characters, the standard question to be expanded is segmented to obtain key words and sentences in the standard question to be expanded and the semantic types corresponding to the key words and sentences;
[0023] According to the semantic type corresponding to each of the key words and phrases in the standard question to be expanded, at least one target expansion question template matching the standard question to be expanded is obtained from a template library, wherein the target expansion question template includes placeholders, the number of the placeholders is the same as the number of the key words and phrases, and the semantic type corresponding to each of the placeholders is respectively the same as the semantic type of each of the key words and phrases.
[0024] In one embodiment, the candidate extension question includes a first candidate extension question, and the step of constructing the candidate extension question according to the standard question to be extended and the target extension question template includes:
[0025] For any of the target extended question templates, the key words are used to replace the placeholders corresponding to the key words in the target extended question template to obtain a first candidate extended question.
[0026] In one embodiment, the candidate extension question further includes a second candidate extension question, and the step of constructing the candidate extension question according to the standard question to be extended and the target extension question template further includes:
[0027] Performing word segmentation on the first candidate expanded question to obtain at least one word or phrase to be expanded of the first candidate expanded question;
[0028] For any word or sentence to be expanded, the word or sentence to be expanded is converted into a domain word vector to obtain a domain word vector corresponding to the word or sentence to be expanded;
[0029] According to the domain word vector corresponding to the to-be-expanded word or sentence, an associated word or sentence associated with the to-be-expanded word or sentence is obtained from a synonym database;
[0030] The associated word or phrase is used to replace the corresponding word or phrase to be expanded in the first candidate expansion question to obtain a second candidate expansion question.
[0031] In one embodiment, determining the recommended extension question corresponding to the standard question to be extended from the candidate extension questions includes:
[0032] Converting the standard question to be expanded into a corresponding semantic word vector, and converting each of the candidate expansion questions into a corresponding semantic word vector;
[0033] According to the semantic word vector corresponding to the standard question to be expanded and the semantic word vector corresponding to each of the candidate expanded questions, a recommended expanded question of the standard question to be expanded is determined from the candidate expanded questions.
[0034] In one embodiment, the character type prediction process is implemented by a prediction network, and the method further includes:
[0035] The prediction network is trained using a preset training set, wherein the preset training set includes multiple sample groups, wherein the sample groups include sample dialogue corpus and a type annotation sequence corresponding to the sample dialogue corpus, wherein the type annotation sequence includes a character type corresponding to each character in the sample dialogue corpus.
[0036] A device for recommending an extended question, the device comprising:
[0037] A generating module, configured to generate a template library according to a historical dialogue corpus, wherein the historical dialogue corpus includes a plurality of historical dialogue corpora, and the template library includes at least one extended question template corresponding to the historical dialogue corpus;
[0038] An acquisition module, used for acquiring a target extension question template matching the standard question to be extended from the template library;
[0039] A construction module, configured to construct a candidate extension question according to the standard question to be extended and the target extension question template;
[0040] The determination module is used to determine the recommended extension question corresponding to the standard question to be extended from the candidate extension questions.
[0041] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0042] Generate a template library according to a historical dialogue corpus, wherein the historical dialogue corpus includes a plurality of historical dialogue corpora, and the template library includes at least one extended question template corresponding to the historical dialogue corpus;
[0043] Acquire a target extension question template matching the standard question to be extended from the template library;
[0044] Constructing a candidate extension question according to the standard question to be extended and the target extension question template;
[0045] A recommended extension question corresponding to the standard question to be extended is determined from the candidate extension questions.
[0046] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:
[0047] Generate a template library according to a historical dialogue corpus, wherein the historical dialogue corpus includes a plurality of historical dialogue corpora, and the template library includes at least one extended question template corresponding to the historical dialogue corpus;
[0048] Acquire a target extension question template matching the standard question to be extended from the template library;
[0049] Constructing a candidate extension question according to the standard question to be extended and the target extension question template;
[0050] A recommended extension question corresponding to the standard question to be extended is determined from the candidate extension questions.
[0051] The above-mentioned method, apparatus, computer device and storage medium for recommending extended questions can generate a template library based on a historical dialogue corpus, wherein the historical dialogue corpus includes multiple historical dialogue corpora, and the template library includes at least one extended question template corresponding to the historical dialogue corpus. A target extended question template that matches the standard question to be expanded is obtained from the template library, and candidate extended questions are constructed based on the standard question to be expanded and the target extended question template, and then a recommended extended question corresponding to the standard question to be expanded is obtained based on the candidate extended questions. The method, apparatus, computer device and storage medium for recommending extended questions provided in the embodiments of the present disclosure are such that the extended question template is generated based on a large amount of historical dialogue corpus, which alleviates the dependence of the process of recommending extended questions on manual annotation, thereby reducing labor costs and greatly improving the accuracy of extended question recommendations. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 A flowchart of a recommended method for extending a question in one embodiment;
[0053] Figure 2 A flowchart of recommended method steps for extending the question in one embodiment;
[0054] Figure 3 A flowchart of recommended method steps for extending the question in one embodiment;
[0055] Figure 4 A flowchart of recommended method steps for extending the question in one embodiment;
[0056] Figure 5 A flowchart of recommended method steps for extending the question in one embodiment;
[0057] Figure 6 A flowchart of recommended method steps for extending the question in one embodiment;
[0058] Figure 7 A schematic diagram of a recommended method for extending a question in one embodiment;
[0059] Figure 8 A structural block diagram of a recommended device for extending the question in one embodiment;
[0060] Fig. 9 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0061] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0062] In one embodiment, Figure 1 As shown, a method for recommending an extended question is provided. This embodiment uses the method applied to a terminal as an example for illustration. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0063] Step 102: Generate a template library based on a historical dialogue corpus, wherein the historical dialogue corpus includes a plurality of historical dialogue corpora, and the template library includes at least one extended question template corresponding to the historical dialogue corpus.
[0064] In the disclosed embodiment, the historical dialogue corpus includes a plurality of historical dialogue corpora, and the historical dialogue corpora may include consultation information of customers in specified scenarios. For example, in a banking business scenario, the historical dialogue corpus may include text information corresponding to voice information when customers consult through voice, or text information entered by customers through online consultation, etc.
[0065] A historical dialogue corpus can be constructed in advance based on the collected historical dialogue corpus, and a template library can be generated based on the historical dialogue corpus. For example, keywords in the historical dialogue corpus can be analyzed, and an extended question template corresponding to the historical dialogue corpus can be generated based on the historical dialogue corpus and the corresponding keywords. Further, a template library can be constructed based on multiple extended question templates generated based on the historical dialogue corpus. Among them, the part corresponding to the keyword in the extended question template is a replaceable part.
[0066] Step 104: Acquire a target extension question template that matches the standard question to be extended from a template library.
[0067] For example, after determining the standard question to be extended, a target extended question template matching the standard question to be extended can be obtained from the template library. Exemplarily, an extended question template whose keywords are consistent with or associated with the keywords of the standard question to be extended can be obtained as the target extended question template matching the standard question to be extended.
[0068] Step 106: construct candidate extension questions according to the standard question to be extended and the target extension question template.
[0069] For example, after obtaining the target extension question template, a candidate extension question can be constructed based on the standard question to be extended and the target extension question template. For example, the keyword part in the target extension question template is replaced with the keyword in the standard question to be extended, or the keyword part in the target extension question template is replaced with the associated words of the keyword in the standard question to be extended, so as to construct a candidate extension question of the standard question to be extended.
[0070] Step 108: Determine a recommended extended question corresponding to the standard question to be extended from the candidate extended questions.
[0071] For example, after constructing candidate extension questions, the recommended extension questions corresponding to the standard question to be extended can be determined from the candidate extension questions. For example, when the number of candidate extension questions is small, all candidate extension questions can be used as recommended extension questions corresponding to the standard question to be extended, or when the number of candidate extension questions is large, the candidate extension question with the highest correlation with the standard question to be extended can be selected as the recommended extension question corresponding to the standard question to be extended.
[0072] The above-mentioned method for recommending extended questions can generate a template library based on a historical dialogue corpus, the historical dialogue corpus includes multiple historical dialogue corpora, and the template library includes at least one extended question template corresponding to the historical dialogue corpus. A target extended question template that matches the standard question to be expanded is obtained from the template library, and a candidate extended question is constructed based on the standard question to be expanded and the target extended question template, and then a recommended extended question corresponding to the standard question to be expanded is obtained based on the candidate extended question. In the method for recommending extended questions provided by the embodiments of the present disclosure, the extended question template is generated based on a large amount of historical dialogue corpus, which alleviates the dependence of the process of recommending extended questions on manual annotation, thereby reducing labor costs and greatly improving the accuracy of extended question recommendations.
[0073] In one embodiment, referring to Figure 2 As shown, the above step 102 may include:
[0074] Step 202, performing character type prediction processing on the historical dialogue corpus in the historical dialogue corpus to obtain a type label sequence corresponding to the historical dialogue corpus, wherein the type label sequence includes the character type corresponding to each character in the historical dialogue corpus;
[0075] Step 204, segmenting the historical dialogue corpus according to the character type corresponding to each character, obtaining the words and sentences to be replaced in the historical dialogue corpus and the semantic types corresponding to the words and sentences to be replaced;
[0076] Step 206, using a placeholder corresponding to the semantic type of the to-be-replaced phrase to replace the to-be-replaced phrase in the historical dialogue corpus to obtain an extended question template;
[0077] Step 208: construct a template library based on the extended question model corresponding to the historical dialogue corpus.
[0078] For example, by predicting the character type of each historical dialogue corpus in the historical dialogue corpus, a type label sequence corresponding to the historical dialogue corpus can be obtained, and the type label sequence includes the character type corresponding to each character in the historical dialogue corpus, that is, the type label sequence consists of the character type corresponding to each character in the historical dialogue corpus.
[0079] In the embodiments of the present disclosure, multiple semantic types can be preset. For example, customer intent can be defined as a five-tuple:<V,N,A,C,E> That is, five semantic types are preset, among which V indicates that the semantic type is action type (for example: query, consultation, etc.), N indicates that the semantic type is business name (for example: debit card, password, etc.), A indicates that the semantic type is business attribute (for example: balance, due date, etc.), C indicates that the semantic type is channel (for example: online banking, mobile banking APP, etc.), and E indicates that the semantic type is abnormal (for example: unavailable, cannot access, etc.).
[0080] Each semantic type has a corresponding character type. For example, for any semantic type, the semantic type has a corresponding first character type and a second character type, wherein the first character type indicates the first character of a word or sentence of the semantic type, and the second character type indicates the middle character of a word or sentence of the semantic type. For example: if the semantic type is V, the corresponding first character type is BV, and the second character type is IV; if the semantic type is N, the corresponding first character type is BN, and the second character type is IN; if the semantic type is A, the corresponding first character type is BA, and the second character type is IA; if the semantic type is C, the corresponding first character type is BC, and the second character type is IC; if the semantic type is E, the corresponding first character type is BE, and the second character type is IE. It should be noted that the character type corresponding to a character that does not correspond to any semantic type is empty (for example, O), and its corresponding semantic type is an empty type.
[0081] Exemplarily, a pre-trained neural network can be used to perform character type prediction processing on historical dialogue corpus to obtain a type label sequence corresponding to the historical dialogue corpus. The presently disclosed embodiment does not impose any limitation on the training process of the neural network. Any training method for a neural network that can be trained to predict a type label sequence of a sentence is applicable to the presently disclosed embodiment.
[0082] When the type tag sequence corresponding to the historical dialogue corpus is obtained, the historical dialogue corpus can be segmented according to the character type corresponding to each character in the historical dialogue corpus in the type tag series to obtain the semantic types corresponding to the words to be replaced and the sentences to be replaced in the historical dialogue corpus. For example, continuous characters whose character types correspond to the same semantic type can be divided into a sentence to be replaced, and the semantic type corresponding to the sentence to be replaced is the semantic type corresponding to the character type.
[0083] After obtaining the words and sentences to be replaced corresponding to the historical dialogue corpus, the corresponding extended question template can be obtained according to the words and sentences to be replaced and the semantic type corresponding to the words and sentences to be replaced. Specifically, any semantic type can have a placeholder corresponding to it, and the words and sentences to be replaced in the historical dialogue corpus can be replaced with the placeholder corresponding to the semantic type of the words and sentences to be replaced, thereby obtaining the extended question template.
[0084] For example, the placeholder corresponding to semantic type V can be preset as "#V", the placeholder corresponding to semantic type N can be preset as "#N", the placeholder corresponding to semantic type A can be preset as "#A", the placeholder corresponding to semantic type C can be preset as "#C", and the placeholder corresponding to semantic type E can be preset as "#E". Assuming the historical dialogue corpus is: I want to check the balance, the type label sequence obtained is<O O B-VI-V O O B-A I-A> , where "BV IV" corresponding to "query" corresponds to semantic type V, and "BA IA" corresponding to "balance" corresponds to semantic type A, then it can be determined that "query" is the phrase to be replaced, with semantic type V, and "balance" is the phrase to be replaced, with semantic type A. Then, the placeholder "#V" corresponding to semantic type V is used to replace "query", and the placeholder "#A" corresponding to semantic type A is used to replace "balance", and the extended question template corresponding to the historical dialogue corpus is generated: I want to #V #A.
[0085] By analogy, the extended question templates corresponding to each historical corresponding corpus in the historical dialogue corpus can be obtained, and after merging and removing duplicates of each extended question template, the corresponding template library can be obtained.
[0086] It should be noted that the above defines customer intent as a five-tuple:<V,N,A,C,E> , that is, presetting five semantic types is only an example in the embodiment of the present disclosure, and is not to be understood as a limitation of the embodiment of the present disclosure. In fact, semantic types can be preset according to actual scene requirements, for example, they can be preset to three semantic types, four semantic types, six semantic types, etc., and the embodiment of the present disclosure does not make specific limitations on this.
[0087] The method for recommending extended questions provided by the embodiments of the present disclosure can analyze the words and sentences to be replaced in the historical dialogue corpus according to the character types corresponding to each character, and then construct the corresponding extended question templates, which can enrich the extended question templates, alleviate the dependence of the process of recommending extended questions on manual annotation, reduce labor costs, and greatly improve the accuracy of extended question recommendations.
[0088] In one embodiment, referring to Figure 3 As shown, the above step 204 may include:
[0089] Step 302, traversing the historical dialogue corpus, and when it is determined that the character type corresponding to the currently traversed character corresponds to the first semantic type, determining the currently traversed character as the first character, and continuing to traverse the next character;
[0090] Step 304, if the traversed second character corresponds to the second semantic type or the empty type, the first character is divided into a phrase to be replaced, and the phrase to be replaced corresponds to the first semantic type; wherein the first semantic type is any semantic type in the semantic type, and the second semantic type is any semantic type in the semantic type except the first semantic type.
[0091] For example, the historical dialogue corpus can be traversed to determine whether the character type corresponding to the currently traversed character is an empty type. If it is an empty type, continue to traverse the next character. Otherwise, determine that the semantic type corresponding to the current character type is the first semantic type, and use the current character as the first character to continue traversing the next character. If the next character is still of the first semantic type, use the next character as the first character and continue to traverse downward until the character type corresponds to the second semantic type (any semantic type except the first semantic type) or the second character of the empty type, and divide the continuous first characters into words to be replaced, and the semantic type corresponding to the words to be replaced is the first semantic type.
[0092] By analogy, until the traversal operation of the historical dialogue corpus is completed, the words and sentences to be replaced in the historical dialogue corpus and the semantic types corresponding to the words and sentences to be replaced can be obtained.
[0093] Taking the above example as an example, traverse the characters in the historical dialogue corpus "I want to check the balance". The character type corresponding to "I" and the semantic type corresponding to "O" are empty characters, so continue to traverse "want". The character type corresponding to "want" and the semantic type corresponding to "O" are empty characters, so continue to traverse "check". The character type corresponding to "check" is "B-V" and the semantic type corresponding to it is V, so take "check" as the first character and continue to traverse "inquire". The character type corresponding to "inquire" is "I-V" and the semantic type corresponding to it is V, so take "inquire" as the first character and continue to traverse "one". The character type corresponding to "one" and the semantic type corresponding to "O" are empty characters, so divide the first characters "check" and "inquire" into the replacement phrase to be replaced with the corresponding semantic type V, "query", and continue to traverse "down". The character type corresponding to "down" and the semantic type corresponding to "O" are empty characters, so continue to traverse "balance". The character type corresponding to "balance" is "B-A" and the semantic type corresponding to it is A, so take "balance" as the first character and continue to traverse "amount". The character type corresponding to "amount" is "I-A" and the semantic type corresponding to it is A, so take "amount" as the first character. At this time, there are no characters to be traversed, so divide the first characters "balance" and "amount" into the replacement phrase to be replaced with the corresponding semantic type A, "balance amount", and end the traversal. That is, the historical dialogue corpus "I want to check the balance" includes the replacement phrase to be replaced with the corresponding semantic type V, "query", and the replacement phrase to be replaced with the corresponding semantic type A, "balance amount".
[0094] The recommended method for extended questions provided by the embodiments of the present disclosure can analyze the replacement phrases to be replaced in the historical dialogue corpus according to the character types corresponding to each character, and then construct corresponding extended question templates, which can enrich the extended question templates, alleviate the dependence on manual annotation in the process of recommending extended questions, reduce labor costs, and greatly improve the accuracy of extended question recommendation.
[0095] In one embodiment, referring to Figure 4 as shown, the above step 104 may include:
[0096] Step 402, perform character type prediction processing on the standard question to be extended to obtain a type label sequence corresponding to the standard question to be extended, where the type label sequence includes the character types corresponding to each character in the standard question to be extended;
[0097] Step 404, perform word segmentation on the standard question to be extended according to the character types corresponding to each character to obtain the keyword phrases in the standard question to be extended and the semantic types corresponding to the keyword phrases;
[0098] Step 406, according to the semantic type corresponding to each key word and phrase in the standard question to be expanded, obtain at least one target expansion question template matching the standard question to be expanded from the template library, the target expansion question template includes placeholders, the number of the placeholders is the same as the number of the key words and phrases, and the semantic type corresponding to each placeholder is respectively the same as the semantic type of each key word and phrase.
[0099] For example, character type prediction processing can be performed on the standard question to be expanded to obtain a type label sequence corresponding to the standard question to be expanded, and the type label sequence can include the character type corresponding to each character in the standard question to be expanded. For the character type prediction processing of the standard question to be expanded, reference can be made to the character type prediction processing for the historical dialogue corpus in the aforementioned embodiment, and the embodiments of the present disclosure will not be repeated here.
[0100] After obtaining the type tag sequence corresponding to the standard question to be expanded, the standard question to be expanded can be segmented according to the character type corresponding to each character of the standard question to be expanded in the type tag sequence to obtain the key words and the semantic types corresponding to the key words in the standard question to be expanded. The segmentation operation for the standard question to be expanded can refer to the process of the segmentation operation for the historical dialogue material in the above real-time example, and the embodiment of the present disclosure will not be repeated here.
[0101] After obtaining the key words and sentences in the standard question to be extended and the semantic types corresponding to the key words and sentences, at least one target extended question template matching the standard question to be extended can be obtained from the template library according to the key words and sentences in the standard question to be extended and the semantic types corresponding to the key words and sentences. Exemplarily, an extended question template whose number of placeholders and semantic types corresponding to the placeholders are completely consistent with the number of key words and sentences in the standard question to be extended and the semantic types corresponding to the key words and sentences can be determined, and the extended question template is used as the target extended question template matching the standard question to be extended.
[0102] For example, assuming that the current standard question to be extended is: I mainly want to consult about the expiration date today, after word segmentation, the key words and phrases obtained include "consultation" corresponding to the semantic type V and "expiration date" corresponding to the semantic type A, then an extended question template with two placeholders, and the placeholders are #V corresponding to the semantic type V and #A corresponding to the semantic type A, can be searched from the template library. Assuming that the extended question template "I want to #V a #A" is found, it can be determined that the extended question template "I want to #V a #A" is the target extended question template of the standard question to be extended.
[0103] The method for recommending extended questions provided in the embodiment of the present disclosure can obtain a target extended question template that matches the key words and phrases in the standard question to be expanded from the template library according to the semantic type of the key words and phrases in the standard question to be expanded. Since the dependence of the process of recommending extended questions on manual annotation can be alleviated, the labor cost can be reduced and the accuracy of the extended question recommendation can be greatly improved. In addition, since the extended question template in the embodiment of the present disclosure is constructed based on the semantic type, the constructed recommended extended questions can not only cover the high-frequency intent expression mode, but also cover the long-tail intent expression mode, which can improve the applicability of the recommended extended questions.
[0104] In one embodiment, the candidate extension question includes a first candidate extension question, and the above step 106 may include:
[0105] For any target extended question template, the key words are used to replace the placeholders corresponding to the key words in the target extended question template to obtain the first candidate extended question.
[0106] For example, after obtaining the target extended question template corresponding to the standard question to be extended, the key words and phrases in the standard question to be extended can be used to replace the placeholders of the same semantic type corresponding to the key words and phrases in the target extended question template, thereby obtaining the first candidate extended question.
[0107] Still taking the above example, the standard question to be extended includes "consultation" corresponding to semantic type V and "expiration date" corresponding to semantic type A, and the target extended question template is "I want to #V about #A". Then, the key phrase "consultation" corresponding to semantic type V can be used to replace "#V" in the target extended question template, and the key phrase "expiration date" corresponding to semantic type A can be used to replace "#A" in the target extended question template, and the first candidate extended question is obtained: I want to consult about the expiration date.
[0108] Similarly, multiple first candidate extension questions can be obtained through each target extension question template, and then recommended extension questions can be obtained based on the first candidate extension questions, for example: the first candidate extension question is used as the recommended extension question, or the first candidate extension question with a higher correlation with the standard question to be extended is selected from the first candidate extension questions as the recommended extension question.
[0109] The method for recommending extended questions provided by the embodiments of the present disclosure can construct an extended question template based on historical dialogue data, and after selecting a target extended question template therefrom, construct a first candidate extended question according to the standard question to be expanded and the target extended question template, and determine a recommended extended question from the first candidate extended question, which can alleviate the dependence of the process of recommending extended questions on manual labeling, reduce labor costs, and greatly improve the accuracy of extended question recommendations.
[0110] In one embodiment, the candidate extension question also includes a second candidate extension question, referring to Figure 5 , the above step 106 may include:
[0111] Step 502, segmenting the first candidate expansion question to obtain at least one word or phrase to be expanded of the first candidate expansion question;
[0112] Step 504: for any word or phrase to be expanded, the word or phrase to be expanded is converted into a domain word vector to obtain a domain word vector corresponding to the word or phrase to be expanded;
[0113] Step 506, according to the domain word vector corresponding to the word to be expanded, obtain the associated words and sentences associated with the word to be expanded from the synonym database;
[0114] Step 508: Use the associated words and phrases to replace the corresponding words and phrases to be expanded in the first candidate expansion question to obtain a second candidate expansion question.
[0115] For example, the first candidate expansion question may be segmented to obtain at least one word to be expanded of the first candidate expansion question. The disclosed embodiment does not specifically limit the word segmentation method, which includes but is not limited to: forward maximum matching method, reverse maximum matching method, bidirectional maximum matching method, etc. After word segmentation, non-preset stop words in the obtained words and sentences may be used as words and sentences to be expanded.
[0116] For example, the historical dialogue corpus in the same field can be used to construct a synonym library (for example, finance and other fields). The domain word vector model can be trained using methods such as CBOW (Continuous Bag-of-Words) or Skip-gram model, so that the domain word vector of each word and sentence in the historical dialogue corpus in the corresponding field can be obtained according to the domain word vector model, and then the corresponding synonym library can be constructed according to the similarity of the domain word vectors of each word and sentence.
[0117] The first candidate expansion question can be traversed, and word segmentation processing can be performed on it to obtain the words and sentences to be expanded of the first candidate expansion question, and after the words and sentences to be expanded are converted into corresponding domain word vectors through the domain word vector model, the domain word vector is used to obtain N synonymous words and sentences with high similarity to the domain word vector from synonyms, and the N synonymous words and sentences are used as associated words and sentences of the words and sentences to be expanded (where N can be a preset value or a value obtained according to a preset ratio, etc.). The N associated words and sentences can be replaced with the corresponding words and sentences to be expanded, thereby obtaining N second candidate expansion questions.
[0118] Similarly, for each word to be expanded in the first candidate expansion question, a corresponding related word can be obtained, and each related word can be used to replace the word to be expanded in the first candidate expansion question to obtain multiple second candidate expansion questions. At this time, the candidate expansion questions can include the first candidate expansion question and the second candidate expansion question, and then the recommended expansion question can be obtained from the candidate expansion questions.
[0119] The method for recommending extended questions provided by the embodiments of the present disclosure trains a domain word vector model that tends to be colloquial through historical dialogue data. A synonym library can be constructed through the domain word vector model, and there is no need to manually maintain the synonym library. The synonym generalization ability of extended questions can be improved, and the synonym library can be used to enrich candidate extended questions, thereby alleviating problems such as a single expression form, poor diversity, and insufficient number of recommendations for recommended extended questions, thereby improving the accuracy of customer intent recognition and the generalization ability of the model.
[0120] In one embodiment, referring to Figure 6 , the above step 108 may include:
[0121] Step 602, converting the standard question to be expanded into a corresponding semantic word vector, and converting each candidate expansion question into a corresponding semantic word vector;
[0122] Step 604: Determine a recommended extension question for the standard question to be extended from the candidate extension questions based on the semantic word vector corresponding to the standard question to be extended and the semantic word vector corresponding to each of the candidate extension questions.
[0123] For example, a language model can be trained using an autoencoder or autoregressive method, and the language model can realize vectorized representation of text information. The language model can be used to convert the standard question to be expanded into a corresponding semantic word vector, and to convert each candidate expansion question into a corresponding semantic word vector, and through a semantic similarity calculation method (for example: cosine similarity algorithm, etc.), the similarity between the semantic word vector corresponding to the standard question to be expanded and the semantic word vector corresponding to each candidate expansion question is calculated respectively, and the recommended expansion question of the standard question to be expanded is determined from the candidate expansion questions according to the similarity. For example: a candidate expansion question with a similarity greater than or equal to a similarity threshold can be selected from the candidate expansion questions as the recommended expansion question of the standard question to be expanded; or, the M candidate expansion questions with the highest similarity can be selected from the candidate expansion questions as the recommended expansion question of the standard question to be expanded, where M is a preset value.
[0124] Furthermore, after obtaining a plurality of recommended extension questions, they may be arranged in descending order according to the magnitude of the similarity corresponding to each recommended extension question, thereby obtaining a recommended extension question list.
[0125] The method for recommending extended questions provided by the embodiment of the present disclosure can obtain a large number of candidate extended questions through extended question templates and a synonym library, and can obtain recommended extended questions from the candidate extended questions, which can improve the accuracy of recommended extended questions and further improve the accuracy of customer intent recognition.
[0126] In one embodiment, the character type prediction process may be implemented by a prediction network, and the above method may further include:
[0127] A preset training set is used to train the prediction network. The preset training set includes multiple sample groups. The sample group includes a sample dialogue corpus and a type annotation sequence corresponding to the sample dialogue corpus. The type annotation sequence includes a character type corresponding to each character in the sample dialogue corpus.
[0128] For example, the intelligent customer service dialogue system often makes corresponding feedback and responses based on the intention of the customer's inquiry. Taking the above example as an example, combined with the manually constructed actions, services, attributes, channels and abnormal situations, the embodiment of this disclosure defines the customer intent as a five-tuple:<V,N,A,C,E> , which corresponds to five semantic types. The semantic types corresponding to V, N, A, C, and E can refer to the above-mentioned embodiment, which will not be described in detail in the embodiment of the present disclosure.
[0129] Sample dialogue corpora can be pre-screened from large-scale historical dialogue corpora. The BIO method can be used to annotate the character types of characters in the sample dialogue corpora (including labels such as B_V, I_V, B_N, I_N, B_A, I_A, B_C, I_C, B_E, I_E, O, etc.) to obtain a type annotation sequence of the sample dialogue corpora. According to each sample dialogue corpus and the type annotation sequence corresponding to the sample dialogue corpus, multiple sample groups can be obtained, and then a preset training set can be constructed based on the multiple sample groups.
[0130] Each sample dialogue data can be input into the prediction network for character type prediction processing to obtain a type label sequence corresponding to the sample dialogue data. According to the type label sequence and type annotation sequence corresponding to the sample dialogue data, the network loss of the prediction network can be determined, and the network parameters of the prediction network can be adjusted according to the network loss until the network loss of the prediction network meets the training requirements (for example, the network loss is less than the loss threshold), and the training of the prediction network is completed to obtain a pre-trained prediction network.
[0131] In the embodiment of the present disclosure, the preset training set can be further split into a training set and a test set. After the prediction model is trained on the training set, the effect of the prediction model is evaluated on the test set, and the prediction model with the best effect on the test set can be selected as the core model for subsequent expansion question recommendation.
[0132] For example: the prediction model is used to perform character type prediction processing on historical dialogue materials to obtain the type label sequence of the historical dialogue materials, and then an extension question template is constructed based on the type label sequence of the historical dialogue materials; or the prediction model is used to perform character type prediction processing on standard questions to be extended to obtain the type label sequence corresponding to the standard questions to be extended, and then a target extension question template is determined based on the type label sequence corresponding to the standard questions to be extended, and a candidate extension question is constructed based on the target extension question template.
[0133] The method for recommending extended questions provided by the embodiments of the present disclosure can realize the construction of extended question templates and the construction of candidate extended questions by using a lightweight prediction network, thereby improving the recommendation speed.
[0134] In order to enable those skilled in the art to better understand the embodiments of the present disclosure, the embodiments of the present disclosure are described below through specific examples.
[0135] Reference Figure 7 As shown, a preset training set can be constructed based on a large-scale historical dialogue corpus, and the preset training set can be used to train the prediction model. The prediction model is used to perform character type prediction processing on the historical dialogue corpus in the historical dialogue corpus, and after obtaining the type label sequence corresponding to each historical dialogue corpus, the words and sentences to be replaced and the semantic types corresponding to the words and sentences to be replaced corresponding to each historical dialogue corpus are determined according to the type label sequence corresponding to each historical dialogue corpus. After the corresponding extended question templates are constructed according to the words and sentences to be replaced and the semantic types corresponding to the words and sentences to be replaced corresponding to each historical dialogue corpus, the extended question templates are merged, deduplicated, and other processes are performed to obtain the corresponding template library.
[0136] The prediction model is used to perform character type prediction processing on the standard question to be extended, and the type label sequence corresponding to the standard question to be extended is obtained, and the key words and sentences of the standard question to be extended and the semantic types corresponding to the key words and sentences are determined according to the type label sequence of the standard question to be extended. The key words and sentences and the semantic types corresponding to the key words and sentences are determined according to the type label sequence of the standard question to be extended, and the target extended question template matching the standard question to be extended is obtained from the template library, and the key words and sentences of the standard question to be extended are used to replace the placeholders in the target extended question template, so as to obtain the first candidate extended question.
[0137] The domain word vector model can be pre-trained, and after the first candidate expansion question is segmented, the words and sentences to be expanded of the first candidate expansion question can be obtained. The domain word vectors corresponding to each word and sentence to be expanded can be obtained according to the domain word vector model, and the associated words and sentences of each word and sentence to be expanded can be determined from the synonym library constructed based on the domain word vector model and the historical dialogue corpus according to the domain word vector corresponding to each word and sentence to be expanded. The associated words and sentences of each word and sentence to be expanded are used to replace the corresponding words and sentences to be expanded in the first candidate expansion question, and the second candidate expansion question can be obtained.
[0138] After using a pre-trained language model to convert the standard question to be expanded into a corresponding semantic word vector, and converting each candidate expansion question (including the first candidate expansion question and the second candidate expansion question) into a corresponding semantic word vector, a recommended standard question for the standard question to be expanded can be obtained from the candidate expansion questions based on the similarity between the semantic word vector corresponding to the standard question to be expanded and the semantic word vector corresponding to each candidate expansion question.
[0139] The method for recommending extended questions provided by the embodiments of the present disclosure builds a template library based on rich historical dialogue corpus, can obtain large-scale recommended extended questions, and can improve the diversity of extended questions. The recommended extended questions obtained based on the extended question templates in the template library can not only cover high-frequency intent expression patterns, but also cover long-tail intent expression patterns, which can improve the applicability of extended questions. In addition, the embodiments of the present disclosure train a domain word vector model based on historical dialogue corpus, and create a synonym library based on the domain word vector model, which can learn synonyms from large-scale domain corpus, without the need for manual maintenance of the synonym library, which can reduce the investment in manually constructing the synonym library and can improve the generalization ability of extended questions. The neural networks involved in the embodiments of the present disclosure are all lightweight networks, so the response speed is fast.
[0140] It should be understood that although Figure 1-7 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1-7 At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.
[0141] In one embodiment, Figure 8 As shown, a device for recommending an extended question is provided, comprising a generating module 802, an acquiring module 804, a constructing module 806 and a determining module 808, wherein:
[0142] A generating module 802 is used to generate a template library according to a historical dialogue corpus, wherein the historical dialogue corpus includes a plurality of historical dialogue corpora, and the template library includes at least one extended question template corresponding to the historical dialogue corpus;
[0143] An acquisition module 804 is used to acquire a target extension question template matching the standard question to be extended from a template library;
[0144] A construction module 806, configured to construct a candidate extension question according to the standard question to be extended and the target extension question template;
[0145] The determination module 808 is used to determine the recommended extension question corresponding to the standard question to be extended from the candidate extension questions.
[0146] The above-mentioned device for recommending extended questions can generate a template library based on a historical dialogue corpus, the historical dialogue corpus includes multiple historical dialogue corpora, and the template library includes at least one extended question template corresponding to the historical dialogue corpus. A target extended question template that matches the standard question to be expanded is obtained from the template library, and a candidate extended question is constructed based on the standard question to be expanded and the target extended question template, and then a recommended extended question corresponding to the standard question to be expanded is obtained based on the candidate extended question. In the device for recommending extended questions provided by the embodiment of the present disclosure, the extended question template is generated based on a large amount of historical dialogue corpus, which alleviates the dependence of the process of recommending extended questions on manual annotation, thereby reducing labor costs and greatly improving the accuracy of extended question recommendations.
[0147] In one embodiment, the generating module 802 is further used to:
[0148] Performing character type prediction processing on the historical dialogue corpus in the historical dialogue corpus to obtain a type label sequence corresponding to the historical dialogue corpus, wherein the type label sequence includes the character type corresponding to each character in the historical dialogue corpus;
[0149] According to the character type corresponding to each character, the historical dialogue corpus is segmented to obtain the words and sentences to be replaced in the historical dialogue corpus and the semantic types corresponding to the words and sentences to be replaced;
[0150] Using the placeholder corresponding to the semantic type of the words and sentences to be replaced, replace the words and sentences to be replaced in the historical dialogue material to obtain an extended question template;
[0151] According to the extended question model corresponding to the historical dialogue corpus, a template library is constructed.
[0152] In one embodiment, the generating module 802 is further used to:
[0153] Traversing the historical dialogue corpus, and when it is determined that the character type corresponding to the currently traversed character corresponds to the first semantic type, determining the currently traversed character as the first character, and continuing to traverse the next character;
[0154] If the traversed second character is found whose character type corresponds to the second semantic type or the empty type, the first character is divided into a phrase to be replaced, and the phrase to be replaced corresponds to the first semantic type;
[0155] The first semantic type is any semantic type in the semantic types, and the second semantic type is any semantic type in the semantic types except the first semantic type.
[0156] In one embodiment, the acquisition module 804 is further used to:
[0157] Perform character type prediction processing on the standard question to be extended to obtain a type label sequence corresponding to the standard question to be extended, wherein the type label sequence includes the character type corresponding to each character in the standard question to be extended;
[0158] According to the character type corresponding to each character, the standard question to be expanded is segmented to obtain the key words and the semantic types corresponding to the key words in the standard question to be expanded;
[0159] According to the semantic type corresponding to each key word and phrase in the standard question to be expanded, at least one target expansion question template matching the standard question to be expanded is obtained from the template library, the target expansion question template includes placeholders, the number of placeholders is the same as the number of key words and phrases, and the semantic type corresponding to each placeholder is the same as the semantic type of each key word and phrase.
[0160] In one embodiment, the candidate extension question includes a first candidate extension question, and the acquisition module 804 is further used to:
[0161] For any target extended question template, the key words are used to replace the placeholders corresponding to the key words in the target extended question template to obtain the first candidate extended question.
[0162] In one embodiment, the candidate extension question further includes a second candidate extension question, and the acquisition module 804 is further used to:
[0163] Performing word segmentation on the first candidate expansion question to obtain at least one to-be-expanded word or phrase of the first candidate expansion question;
[0164] For any word or sentence to be expanded, the word or sentence to be expanded is converted into a domain word vector to obtain the domain word vector corresponding to the word or sentence to be expanded;
[0165] According to the domain word vector corresponding to the to-be-expanded word or sentence, the associated words or sentences associated with the to-be-expanded word or sentence are obtained from the synonym database;
[0166] The associated words and sentences are used to replace the corresponding words and sentences to be expanded in the first candidate expansion question to obtain the second candidate expansion question.
[0167] In one embodiment, the determination module 808 is further configured to:
[0168] Convert the standard question to be expanded into the corresponding semantic word vector, and convert each candidate expansion question into the corresponding semantic word vector;
[0169] According to the semantic word vector corresponding to the standard question to be expanded and the semantic word vector corresponding to each candidate expansion question, a recommended expansion question of the standard question to be expanded is determined from the candidate expansion questions.
[0170] In one embodiment, the character type prediction process is implemented by a prediction network, and the device further includes:
[0171] A training module is used to train the prediction network using a preset training set, wherein the preset training set includes multiple sample groups, wherein the sample groups include sample dialogue corpus and a type annotation sequence corresponding to the sample dialogue corpus, wherein the type annotation sequence includes a character type corresponding to each character in the sample dialogue corpus.
[0172] The specific definition of the device for recommending extended questions can be found in the definition of the method for recommending extended questions in the above text, and will not be repeated here. Each module in the above-mentioned device for recommending extended questions can be implemented in whole or in part by software, hardware, and a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0173] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Fig. 9 As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a recommended method for extending the question is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a key, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0174] Those skilled in the art will understand that Fig. 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0175] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:
[0176] A template library is generated according to a historical dialogue corpus, wherein the historical dialogue corpus includes a plurality of historical dialogue corpora, and the template library includes at least one extended question template corresponding to the historical dialogue corpus; a target extended question template matching a standard question to be extended is acquired from the template library; a candidate extended question is constructed according to the standard question to be extended and the target extended question template; and a recommended extended question corresponding to the standard question to be extended is determined from the candidate extended questions.
[0177] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0178] Perform character type prediction processing on historical dialogue corpus in a historical dialogue corpus to obtain a type label sequence corresponding to the historical dialogue corpus, wherein the type label sequence includes the character type corresponding to each character in the historical dialogue corpus; perform word segmentation on the historical dialogue corpus according to the character type corresponding to each character to obtain words and sentences to be replaced in the historical dialogue corpus and semantic types corresponding to the words and sentences to be replaced; use placeholders corresponding to the semantic types of the words and sentences to be replaced in the historical dialogue corpus to replace the words and sentences to be replaced in the historical dialogue corpus to obtain extended question templates; and construct a template library according to the extended question model corresponding to the historical dialogue corpus.
[0179] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0180] Traverse the historical dialogue corpus, and when it is determined that the character type corresponding to the currently traversed character corresponds to the first semantic type, determine that the currently traversed character is the first character, and continue to traverse the next character; if the traversed second character corresponds to the second semantic type or the empty type, classify the first character into a phrase to be replaced, and the phrase to be replaced corresponds to the first semantic type; wherein the first semantic type is any semantic type among the semantic types, and the second semantic type is any semantic type among the semantic types except the first semantic type.
[0181] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0182] A character type prediction process is performed on the standard question to be expanded to obtain a type label sequence corresponding to the standard question to be expanded, wherein the type label sequence includes the character type corresponding to each character in the standard question to be expanded; according to the character type corresponding to each character, the standard question to be expanded is segmented to obtain key words and phrases in the standard question to be expanded and the semantic types corresponding to the key words and phrases; according to the semantic types corresponding to each key word and phrase in the standard question to be expanded, at least one target expansion question template matching the standard question to be expanded is obtained from a template library, wherein the target expansion question template includes placeholders, the number of the placeholders is the same as the number of the key words and phrases, and the semantic types corresponding to each placeholder are respectively the same as the semantic types of each key word and phrase.
[0183] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0184] For any of the target extended question templates, the key words are used to replace the placeholders corresponding to the key words in the target extended question template to obtain a first candidate extended question.
[0185] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0186] The first candidate expansion question is segmented to obtain at least one word or phrase to be expanded of the first candidate expansion question; for any word or phrase to be expanded, the word or phrase to be expanded is converted into a domain word vector to obtain a domain word vector corresponding to the word or phrase to be expanded; according to the domain word vector corresponding to the word or phrase to be expanded, an associated word or phrase associated with the word or phrase to be expanded is obtained from a synonym library; the associated word or phrase is used to replace the corresponding word or phrase to be expanded in the first candidate expansion question to obtain a second candidate expansion question.
[0187] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0188] The standard question to be expanded is converted into a corresponding semantic word vector, and each of the candidate expansion questions is converted into a corresponding semantic word vector; based on the semantic word vector corresponding to the standard question to be expanded and the semantic word vector corresponding to each of the candidate expansion questions, a recommended expansion question of the standard question to be expanded is determined from the candidate expansion questions.
[0189] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0190] The prediction network is trained using a preset training set, wherein the preset training set includes multiple sample groups, wherein the sample groups include sample dialogue corpus and a type annotation sequence corresponding to the sample dialogue corpus, wherein the type annotation sequence includes a character type corresponding to each character in the sample dialogue corpus.
[0191] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0192] A template library is generated according to a historical dialogue corpus, wherein the historical dialogue corpus includes a plurality of historical dialogue corpora, and the template library includes at least one extended question template corresponding to the historical dialogue corpus; a target extended question template matching a standard question to be extended is acquired from the template library; a candidate extended question is constructed according to the standard question to be extended and the target extended question template; and a recommended extended question corresponding to the standard question to be extended is determined from the candidate extended questions.
[0193] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0194] Perform character type prediction processing on historical dialogue corpus in a historical dialogue corpus to obtain a type label sequence corresponding to the historical dialogue corpus, wherein the type label sequence includes the character type corresponding to each character in the historical dialogue corpus; perform word segmentation on the historical dialogue corpus according to the character type corresponding to each character to obtain words and sentences to be replaced in the historical dialogue corpus and semantic types corresponding to the words and sentences to be replaced; use placeholders corresponding to the semantic types of the words and sentences to be replaced in the historical dialogue corpus to replace the words and sentences to be replaced in the historical dialogue corpus to obtain extended question templates; and construct a template library according to the extended question model corresponding to the historical dialogue corpus.
[0195] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0196] Traverse the historical dialogue corpus, and when it is determined that the character type corresponding to the currently traversed character corresponds to the first semantic type, determine that the currently traversed character is the first character, and continue to traverse the next character; if the traversed second character corresponds to the second semantic type or the empty type, classify the first character into a phrase to be replaced, and the phrase to be replaced corresponds to the first semantic type; wherein the first semantic type is any semantic type among the semantic types, and the second semantic type is any semantic type among the semantic types except the first semantic type.
[0197] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0198] A character type prediction process is performed on the standard question to be expanded to obtain a type label sequence corresponding to the standard question to be expanded, wherein the type label sequence includes the character type corresponding to each character in the standard question to be expanded; according to the character type corresponding to each character, the standard question to be expanded is segmented to obtain key words and phrases in the standard question to be expanded and the semantic types corresponding to the key words and phrases; according to the semantic types corresponding to each key word and phrase in the standard question to be expanded, at least one target expansion question template matching the standard question to be expanded is obtained from a template library, wherein the target expansion question template includes placeholders, the number of the placeholders is the same as the number of the key words and phrases, and the semantic types corresponding to each placeholder are respectively the same as the semantic types of each key word and phrase.
[0199] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0200] For any of the target extended question templates, the key words are used to replace the placeholders corresponding to the key words in the target extended question template to obtain a first candidate extended question.
[0201] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0202] The first candidate expansion question is segmented to obtain at least one word or phrase to be expanded of the first candidate expansion question; for any word or phrase to be expanded, the word or phrase to be expanded is converted into a domain word vector to obtain a domain word vector corresponding to the word or phrase to be expanded; according to the domain word vector corresponding to the word or phrase to be expanded, an associated word or phrase associated with the word or phrase to be expanded is obtained from a synonym library; the associated word or phrase is used to replace the corresponding word or phrase to be expanded in the first candidate expansion question to obtain a second candidate expansion question.
[0203] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0204] The standard question to be expanded is converted into a corresponding semantic word vector, and each of the candidate expansion questions is converted into a corresponding semantic word vector; based on the semantic word vector corresponding to the standard question to be expanded and the semantic word vector corresponding to each of the candidate expansion questions, a recommended expansion question of the standard question to be expanded is determined from the candidate expansion questions.
[0205] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0206] The prediction network is trained using a preset training set, wherein the preset training set includes multiple sample groups, wherein the sample groups include sample dialogue corpus and a type annotation sequence corresponding to the sample dialogue corpus, wherein the type annotation sequence includes a character type corresponding to each character in the sample dialogue corpus.
[0207] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0208] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0209] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A method for recommending an extended question, characterized in that: The method comprises: Generate a template library according to a historical dialogue corpus, wherein the historical dialogue corpus includes a plurality of historical dialogue corpora, and the template library includes at least one extended question template corresponding to the historical dialogue corpus; Performing character type prediction processing on the standard question to be expanded to obtain a type label sequence corresponding to the standard question to be expanded, wherein the type label sequence includes the character type corresponding to each character in the standard question to be expanded; According to the character type corresponding to each of the characters, the standard question to be expanded is segmented to obtain key words and sentences in the standard question to be expanded and the semantic types corresponding to the key words and sentences; According to the semantic type corresponding to each of the key words and phrases in the standard question to be expanded, obtaining at least one target expansion question template matching the standard question to be expanded from a template library, wherein the target expansion question template includes placeholders, the number of the placeholders is the same as the number of the key words and phrases, and the semantic type corresponding to each of the placeholders is respectively the same as the semantic type of each of the key words and phrases; Constructing a candidate extension question according to the standard question to be extended and the target extension question template; Determining, from the candidate extended questions, a recommended extended question corresponding to the standard question to be extended; The step of constructing a candidate extension question according to the standard question to be extended and the target extension question template includes: The key words in the target extended question template are partially replaced with key words and phrases in the standard question to be extended, or the key words in the target extended question template are partially replaced with associated words and phrases of the key words and phrases in the standard question to be extended, so as to construct a candidate extended question of the standard question to be extended.
2. The method according to claim 1, characterized in that The step of generating a template library based on the historical dialogue corpus includes: Performing character type prediction processing on historical dialogue corpus in the historical dialogue corpus to obtain a type label sequence corresponding to the historical dialogue corpus, wherein the type label sequence includes the character type corresponding to each character in the historical dialogue corpus; According to the character types corresponding to the characters, the historical dialogue corpus is segmented to obtain the words and sentences to be replaced in the historical dialogue corpus and the semantic types corresponding to the words and sentences to be replaced; Using a placeholder corresponding to the semantic type of the to-be-replaced phrase to replace the to-be-replaced phrase in the historical dialogue corpus to obtain an extended question template; A template library is constructed based on the extended question template corresponding to the historical dialogue corpus.
3. The method according to claim 2, characterized in that The segmenting of the historical dialogue corpus according to the character type corresponding to each character to obtain the words and sentences to be replaced in the historical dialogue corpus and the semantic types corresponding to the words and sentences to be replaced includes: Traversing the historical dialogue corpus, and when it is determined that the character type corresponding to the currently traversed character corresponds to the first semantic type, determining the currently traversed character as the first character, and continuing to traverse the next character; If a second character whose character type corresponds to the second semantic type or the empty type is traversed, the first character is divided into a phrase to be replaced, and the phrase to be replaced corresponds to the first semantic type; The first semantic type is any semantic type among the semantic types, and the second semantic type is any semantic type among the semantic types except the first semantic type.
4. The method according to claim 1, characterized in that: The candidate extension question includes a first candidate extension question, and constructing the candidate extension question according to the standard question to be extended and the target extension question template includes: For any of the target extended question templates, the key words are used to replace the placeholders corresponding to the key words in the target extended question template to obtain a first candidate extended question.
5. The method according to claim 4, characterized in that The candidate extension question also includes a second candidate extension question, and the step of constructing the candidate extension question according to the standard question to be extended and the target extension question template further includes: Performing word segmentation on the first candidate expanded question to obtain at least one word or phrase to be expanded of the first candidate expanded question; For any word or sentence to be expanded, the word or sentence to be expanded is converted into a domain word vector to obtain a domain word vector corresponding to the word or sentence to be expanded; According to the domain word vector corresponding to the to-be-expanded word or sentence, an associated word or sentence associated with the to-be-expanded word or sentence is obtained from a synonym database; The associated word or phrase is used to replace the corresponding word or phrase to be expanded in the first candidate expansion question to obtain a second candidate expansion question.
6. The method according to any one of claims 1 to 5, characterized in that: The step of determining, from the candidate extended questions, a recommended extended question corresponding to the standard question to be extended comprises: Converting the standard question to be expanded into a corresponding semantic word vector, and converting each of the candidate expansion questions into a corresponding semantic word vector; According to the semantic word vector corresponding to the standard question to be expanded and the semantic word vector corresponding to each of the candidate expanded questions, a recommended expanded question of the standard question to be expanded is determined from the candidate expanded questions.
7. The method according to claim 1, characterized in that The character type prediction process is implemented by a prediction network, and the method further comprises: The prediction network is trained using a preset training set, wherein the preset training set includes multiple sample groups, wherein the sample groups include sample dialogue corpus and a type annotation sequence corresponding to the sample dialogue corpus, wherein the type annotation sequence includes a character type corresponding to each character in the sample dialogue corpus.
8. A device for recommending an extended question, characterized in that: The device comprises: A generating module, configured to generate a template library according to a historical dialogue corpus, wherein the historical dialogue corpus includes a plurality of historical dialogue corpora, and the template library includes at least one extended question template corresponding to the historical dialogue corpus; An acquisition module is used to perform character type prediction processing on the standard question to be expanded, obtain a type label sequence corresponding to the standard question to be expanded, and the type label sequence includes the character type corresponding to each character in the standard question to be expanded; according to the character type corresponding to each character, the standard question to be expanded is segmented to obtain key words and phrases in the standard question to be expanded and the semantic types corresponding to the key words and phrases; according to the semantic types corresponding to each key word and phrase in the standard question to be expanded, at least one target expansion question template matching the standard question to be expanded is acquired from a template library, the target expansion question template includes placeholders, the number of the placeholders is the same as the number of the key words and phrases, and the semantic types corresponding to each placeholder are respectively the same as the semantic types of each key word and phrase; A construction module, configured to construct a candidate extension question according to the standard question to be extended and the target extension question template; A determination module, configured to determine, from the candidate extension questions, a recommended extension question corresponding to the standard question to be extended; The step of constructing a candidate extension question according to the standard question to be extended and the target extension question template includes: The key words in the target extended question template are partially replaced with key words and phrases in the standard question to be extended, or the key words in the target extended question template are partially replaced with associated words and phrases of the key words and phrases in the standard question to be extended, so as to construct a candidate extended question of the standard question to be extended.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Travel knowledge graph building method and device and travel question answering method and device
CN107729493A
Question generation method and device
CN111222309A