Intelligent knowledge base expansion method, device, electronic device and storage medium
By obtaining the user consultation text prefix and using the GPT-2 model to generate a candidate set of common consultation questions, the problem of similar FAQ knowledge points in the intelligent knowledge base is solved, and a higher user consultation question coverage and more accurate question-and-answer service are achieved.
Patent Information
- Application Number
- CN202111501206.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-09
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2041-12-09
AI Technical Summary
When the prior art expands similar questions about the FAQ knowledge points of the intelligent knowledge base, it cannot effectively cover the text of the user's common consultation question, and it is not accurate enough.
By obtaining the user consultation text prefix, inputting the trained GPT-2 model for text generation, obtaining a candidate set of common consultation questions, and comparing it with the text of the problem in the intelligent knowledge base, and adding text with the similarity higher than the preset threshold to the intelligent knowledge base.
The coverage rate of user consultation questions in the FAQ knowledge points of the intelligent knowledge base has been improved, and the accuracy and service quality of the intelligent customer service robot question and answer are enhanced.
Smart Images

Figure CN114328861B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, device, electronic device and storage medium for expanding an intelligent knowledge base. Background Art
[0002] Frequently Asked Questions (FAQ) is the main means of providing online help on the current network. With the rapid development of artificial intelligence, intelligent customer service robots have become relatively common auxiliary tools in our lives. During the use of intelligent customer service robots, in order to improve the accuracy of the answers of intelligent customer service robots, the user consultation questions of the FAQ knowledge points in the intelligent knowledge base are usually expanded.
[0003] Currently, similar questions are mainly expanded by means of synonym replacement. However, the similar questions expanded from the FAQ knowledge points of the intelligent knowledge base may not be the common consultation question texts of users and may not necessarily cover the user consultation behaviors. Summary of the Invention
[0004] Embodiments of the present application provide a method, device, electronic device and storage medium for expanding an intelligent knowledge base, which are used to improve the coverage rate of user consultation questions of the FAQ knowledge points in the intelligent knowledge base.
[0005] In a first aspect, a method for expanding an intelligent knowledge base is provided. The method includes:
[0006] Obtain a prefix of a user consultation text; wherein, the prefix of the user consultation text is the first character in the user consultation text or a combination of consecutive multiple characters starting from the first character;
[0007] Input the prefix of the user consultation text into a trained text generation model for text generation to obtain a candidate set of common consultation questions; wherein, the text generation model is obtained by training a GPT-2 model, and the embedding layer parameters of the GPT-2 model include the embedding layer parameters corresponding to the pseudo-prompt text and the embedding layer parameters corresponding to the user consultation question text. The pseudo-prompt text is text without an actual task and is used to obtain prompt information during the text generation process. During the training of the GPT-2 model, the embedding layer parameters corresponding to the user consultation question text are masked;
[0008] Compare the texts in the candidate set of common consultation questions with the texts of the questions in the intelligent knowledge base to obtain multiple similarity degrees;
[0009] Add the texts in the candidate set of common consultation questions with a similarity degree higher than a first preset threshold to the intelligent knowledge base.
[0010] Optionally, the text generation model is obtained in the following manner:
[0011] Obtain the text of user consultation questions in a preset field;
[0012] Input the pseudo-prompt text and the text of the user consultation question into the embedding layer of the GPT-2 model for encoding to obtain the prompt encoding information corresponding to the pseudo-prompt text and the text encoding information corresponding to the text of the user consultation question, so that the GPT-2 model generates a prediction text based on the prompt encoding information and the text encoding information; wherein, the prompt encoding information is used to prompt the character to be generated.
[0013] Based on the prediction text, optimize the embedding layer parameters corresponding to the pseudo-prompt text in the embedding layer of the GPT-2 model through a loss function until the similarity between the prediction text generated by the GPT-2 model and the text of the user consultation question is greater than a second preset threshold, and then stop optimizing the embedding layer parameters corresponding to the pseudo-prompt text to obtain the text generation model.
[0014] Optionally, the GPT-2 model generates a prediction text based on the prompt encoding information and the text encoding information, including:
[0015] Generate each character in sequence according to the standard grammar order based on the prompt encoding information and the text encoding information; wherein, the new character is generated based on the generated character sequence, and after the new character is generated, the new character is added to the end of the generated character sequence.
[0016] Optionally, the input of the pseudo-prompt text and the text of the user consultation question into the embedding layer of the GPT-2 model for encoding includes:
[0017] Preprocess the text of the user consultation question to obtain a target text; wherein, the preprocessing includes deleting the text containing sensitive vocabulary, incomplete text, picture URLs, special characters, converting the complex writing to the simplified writing, and converting the English uppercase to the lowercase. The sensitive vocabulary is the vocabulary contained in a preset sensitive database, and the incomplete text is the text whose grammar is different from the standard grammar.
[0018] Input the pseudo-prompt text and the target text into the embedding layer of the GPT-2 model for encoding.
[0019] Optionally, the obtaining of the prefix of the user consultation text includes:
[0020] Crawl the user consultation text data;
[0021] Statistically analyze the text prefixes in the user consultation text data;
[0022] Determine the number of occurrences of each text prefix;
[0023] Sort the text prefixes according to the number of occurrences; among them, the text prefix with the most occurrences is the text prefix with the highest priority;
[0024] Obtain the user consultation text prefix for text generation according to the priority order.
[0025] Optionally, the similarity comparison between the text in the common consultation question candidate set and the text of the questions in the intelligent knowledge base includes:
[0026] Select a preset number of common consultation questions from the common question candidate set through a beam search algorithm;
[0027] Perform a similarity comparison between the text of the preset number of common consultation questions and the text of the questions in the intelligent knowledge base.
[0028] Optionally, the similarity comparison between the text of the preset number of common consultation questions and the text of the questions in the intelligent knowledge base includes:
[0029] Convert the text of the preset number of common consultation questions and the text of the questions in the intelligent knowledge base into word vector sequences with a fixed length respectively;
[0030] Obtain the first semantic feature vector sequence and the second semantic feature vector sequence corresponding to the first word vector sequence and the second word vector sequence respectively through a deep learning algorithm; where, the first word vector sequence is the word vector sequence corresponding to the first text in the preset number of common consultation questions, the second word vector sequence is the word vector sequence corresponding to the second text in the intelligent knowledge base, and the lengths of the first word vector sequence and the second word vector sequence are the same;
[0031] Determine the first global feature vector sequence and the second global feature vector sequence corresponding to the first semantic feature vector and the second semantic feature vector sequence according to the attention mechanism;
[0032] Calculate the similarity between each global feature vector in the first global feature vector sequence and each global feature vector in the second global feature vector to obtain a similarity matrix;
[0033] Determine the similarity between each text in the first text and each text in the second text according to the similarity matrix.
[0034] In a second aspect, there is provided an intelligent knowledge base expansion device, and the device includes:
[0035] A processing module for obtaining the prefix of the user's consultation text; wherein, the prefix of the user's consultation text is the first character in the user's consultation text or a combination of consecutive multiple characters starting from the first character.
[0036] The processing module is further configured to input the prefix of the user's consultation text into a trained text generation model for text generation to obtain a candidate set of common consultation questions; wherein, the text generation model is obtained by training a GPT-2 model, and the embedding layer parameters of the GPT-2 model include the embedding layer parameters corresponding to the pseudo-prompt text and the embedding layer parameters corresponding to the user's consultation question text. The pseudo-prompt text is text without an actual task and is used to obtain prompt information during text generation. During the training of the GPT-2 model, the embedding layer parameters corresponding to the user's consultation question text are masked.
[0037] The processing module is further configured to compare the similarity between the text in the candidate set of common consultation questions and the text of the questions in the intelligent knowledge base to obtain multiple similarities.
[0038] The processing module is further configured to add the text in the candidate set of common consultation questions with a similarity higher than a first preset threshold to the intelligent knowledge base.
[0039] Optionally, the intelligent knowledge base expansion device further includes a model training module for:
[0040] Obtaining the user's consultation question text in a preset domain.
[0041] Inputting the pseudo-prompt text and the user's consultation question text into the embedding layer of the GPT-2 model for encoding to obtain the prompt encoding information corresponding to the pseudo-prompt text and the text encoding information corresponding to the user's consultation question text, so that the GPT-2 model generates a prediction text based on the prompt encoding information and the text encoding information; wherein, the prompt encoding information is used to prompt the characters to be generated.
[0042] Based on the prediction text, optimizing the embedding layer parameters corresponding to the pseudo-prompt text in the embedding layer of the GPT-2 model through a loss function until the similarity between the prediction text generated by the GPT-2 model and the user's consultation question text is greater than a second preset threshold, and then stopping optimizing the embedding layer parameters corresponding to the pseudo-prompt text to obtain the text generation model.
[0043] Optionally, the model training module is specifically configured to:
[0044] Generate each character in sequence according to the standard grammar order based on the prompt encoding information and the text encoding information; wherein, the new character is generated based on the generated character sequence, and after the new character is generated, the new character is added to the end of the generated character sequence.
[0045] Optionally, the processing module is specifically configured to:
[0046] Preprocess the user's consultation question text to obtain a target text; wherein, the preprocessing includes deleting text containing sensitive words, incomplete text, picture URLs, special characters, converting complex writing to simplified writing, and converting English uppercase to lowercase. The sensitive words are the words included in a preset sensitive database, and the incomplete text is text with a grammar different from the standard grammar.
[0047] Input the pseudo-prompt text and the target text into the embedding layer of the GPT-2 model for encoding.
[0048] Optionally, the processing module is specifically configured to:
[0049] Crawl user consultation text data;
[0050] Count the text prefixes in the user consultation text data;
[0051] Determine the number of occurrences of each text prefix;
[0052] Sort the text prefixes according to the number of occurrences for priority; wherein, the text prefix with the most occurrences is the text prefix with the highest priority;
[0053] Obtain the text prefix of the user consultation text for text generation according to the priority order.
[0054] Optionally, the processing module is specifically configured to:
[0055] Select a preset number of common consultation questions from the candidate set of common questions through the beam search algorithm;
[0056] Compare the text of the preset number of common consultation questions with the text of the questions in the intelligent knowledge base for similarity.
[0057] Optionally, the processing module is specifically configured to:
[0058] Convert the text of the preset number of common consultation questions and the text of the questions in the intelligent knowledge base into word vector sequences with a fixed length respectively;
[0059] Obtain the first semantic feature vector sequence and the second semantic feature vector sequence corresponding to the first word vector sequence and the second word vector sequence respectively through a deep learning algorithm; wherein, the first word vector sequence is the word vector sequence corresponding to the first text in the preset number of common consultation question texts, the second word vector sequence is the word vector sequence corresponding to the second text in the intelligent knowledge base, and the lengths of the first word vector sequence and the second word vector sequence are the same;
[0060] Determine the first global feature vector sequence and the second global feature vector sequence corresponding to the first semantic feature vector and the second semantic feature vector sequence according to the attention mechanism;
[0061] Calculate the similarity between each global feature vector in the first global feature vector sequence and each global feature vector in the second global feature vector to obtain a similarity matrix;
[0062] Determine the similarity between each text in the first text and each text in the second text according to the similarity matrix.
[0063] In a third aspect, an electronic device is provided, and the electronic device includes:
[0064] A memory for storing program instructions;
[0065] A processor for calling the program instructions stored in the memory and executing the steps included in any of the methods described in the first aspect according to the obtained program instructions.
[0066] In a fourth aspect, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer-executable instructions for causing a computer to execute the steps included in any of the methods described in the first aspect.
[0067] In a fifth aspect, a computer program product containing instructions is provided, and when the computer program product runs on a computer, it causes the computer to execute the intelligent knowledge base expansion method described in the above various possible implementation manners.
[0068] In an embodiment of the present application, the prefix of the user's consultation text is obtained, and the prefix of the user's consultation text is input into a trained text generation model for text generation to obtain a candidate set of common consultation questions. The text generation model is obtained by training a GPT-2 model. The embedding layer parameters of the GPT-2 model include the embedding layer parameters corresponding to the pseudo-prompt text and the embedding layer parameters corresponding to the user's consultation question text. The pseudo-prompt text is text without an actual task and is used to obtain prompt information during the text generation process. During the training of the GPT-2 model, the embedding layer parameters corresponding to the user's consultation question text are masked; the texts in the candidate set of common consultation questions are compared with the texts of the questions in the intelligent knowledge base to obtain multiple similarities, and the texts in the candidate set of long sword consultation questions with similarities higher than the first preset threshold are added to the intelligent knowledge base.
[0069] That is to say, in the present application, the embedding layer parameters corresponding to the user's consultation question text (i.e., the text position parameters corresponding to the user's consultation question text) in the embedding layer of the GPT-2 model are masked, and the embedding layer parameters corresponding to the pseudo-prompt text are optimized and updated, so that the embedding layer of the GPT-2 model can not only generate prompt information but also encode ordinary text, so that the GPT-2 model can automatically generate text according to the prompt information. After the text is generated, the generated text is compared with the text in the intelligent knowledge base, and the text with a high similarity is added to the intelligent knowledge base. That is, the text generation task of common consultation questions in special scenarios is reconstructed by using the automatic construction of prompt information technology (the pseudo-prompt text generates prompt information through the embedding layer) and text generation technology (generating text according to the prompt information), and the FAQ corpus of the intelligent knowledge base is improved by simulating the user's consultation prefix to generate question texts, effectively covering the online user's consultation behavior, thereby improving the coverage rate of the user's consultation questions for the FAQ knowledge points in the intelligent knowledge base. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application.
[0071] Figure 1 It is a flowchart of a method for expanding an intelligent knowledge base provided by an embodiment of the present application;
[0072] Figure 2 It is a structural block diagram of a device for expanding an intelligent knowledge base provided by an embodiment of the present application;
[0073] Figure 3 It is a schematic structural diagram of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] To make the objectives, technical solutions, and advantages of this application more clear and understandable, the following will, in conjunction with the accompanying drawings in the embodiments of this application, clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the described embodiments are only a part rather than all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without making creative efforts belong to the scope of protection of this application. Without conflict, the embodiments in this application and the features in the embodiments may be arbitrarily combined with each other. And although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0075] The terms "first" and "second" in the specification, claims, and above-mentioned accompanying drawings of this application are used to distinguish different objects, rather than to describe a specific order. In addition, the term "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices. "Multiple" in this application may mean at least two, for example, it may be two, three, or more, and there is no limitation in the embodiments of this application.
[0076] In addition, the term "and / or" in this document is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after without special explanation.
[0077] The following introduces the intelligent knowledge base expansion method provided in the embodiments of this application in conjunction with the drawings in the specification. Please refer to Figure 1 As shown, the process of the intelligent knowledge base expansion method in the embodiments of this application is described as follows:
[0078] Step 101: Obtain the prefix of the user's consultation text;
[0079] Among them, the user consultation text prefix (hereinafter referred to as the text prefix) is the first character of the user consultation text or a combination of consecutive multiple characters starting from the first character. The number of characters included in the text prefix can be set according to requirements. For example, if the user consultation text is "I want to know about WeSure Million Medical Insurance", if the number of characters included in the set text prefix is 1, the text prefix corresponding to this user consultation text is "I", if the number of characters included in the set text prefix is 2, the text prefix corresponding to this user consultation text is "I want", if the number of characters included in the set text prefix is 4, the text prefix corresponding to this user consultation text is "I want to know". In the embodiments of the present application, the text prefix can be obtained by crawling user consultation text data and counting the text prefixes in the user consultation text data. For example, it can be counted according to the character length of the text prefix (i.e., the number of characters included), and after the counting is completed, the number of occurrences of each text prefix is determined, and different text prefixes are sorted according to the number of occurrences. Among them, the text prefix with the highest priority has the most occurrences, and the user consultation text prefix for text generation is obtained according to this priority order. Among them, when a text prefix has a large number of occurrences, it indicates that the problem corresponding to this text prefix is a problem that users are interested in. When a text prefix has a small number of occurrences, it indicates that the problem corresponding to this text prefix is a problem that users are less interested in. Therefore, the text prefix with the most occurrences can be used as the user consultation text prefix for text generation first. In a possible implementation manner, the text prefixes with the number of occurrences less than the preset threshold can also be deleted, and text generation is not performed for this text prefix.
[0080] Step 102: Input the user consultation text prefix into the trained text generation model for text generation to obtain a candidate set of common consultation questions;
[0081] Among them, since this text generation model is obtained by training the GPT-2 model, before inputting the user consultation text prefix into this text generation model for text generation, it is also necessary to train the GPT-2 model so that the text generated by the text generation model is close to the text used for model training.
[0082] The training process of the GPT-2 model will be described below.
[0083] Step S10: Obtain the user consultation question text in the preset domain;
[0084] Step S11: Preprocess the user consultation question text to obtain the target text;
[0085] Among them, preprocessing the text of the user's consultation question includes deleting the text containing sensitive words, incomplete text, picture URLs, special characters, and converting the traditional writing to simplified writing, converting the English capital letters to lowercase, etc., so as to make the format of the final target text unified. Among them, the sensitive words are the words included in the preset sensitive database, and the incomplete text is the text whose grammar is different from the standard grammar. In a possible implementation manner, after obtaining the text of the user's consultation question, the grammar of the text of the user's consultation question can be detected to determine whether the text of the user's consultation question is the same as the standard grammar. If it is not the same, the model is not trained with the text of the user's consultation question. At this time, the text of the user's consultation question needs to be deleted. It should be noted that when deleting the text containing sensitive words, the entire text can be deleted, or only the sensitive words can be deleted, so as to ensure the diversity of the text of the user's consultation question for model training as much as possible.
[0086] Step S12: Train the GPT-2 model with the pseudo prompt text and the target text.
[0087] Among them, the embedding layer parameters of the GPT-2 model include the embedding layer parameters corresponding to the pseudo prompt text (the text position parameters corresponding to the pseudo prompt text) and the embedding layer parameters corresponding to the text of the user's consultation question (the text position parameters corresponding to the text of the user's consultation question). In the embodiment of the present application, the embedding layer parameters corresponding to the target text (the text of the user's consultation question) are masked (masked), and the pseudo prompt text and the text of the user's consultation question are input into the embedding layer of the GPT-2 model for encoding to obtain the prompt encoding information corresponding to the pseudo prompt text and the text encoding information corresponding to the text of the user's consultation question, and a prediction text is generated based on the prompt encoding information and the text encoding information; among them, the pseudo prompt text is the text without an actual task, which is used to obtain the prompt information in the text generation process (that is, the prompt encoding information formed after encoding by the embedding layer), and the prompt encoding information is used to prompt the characters to be generated.
[0088] In a possible implementation manner, the GPT-2 model generates each character in turn according to the standard grammar order (for example, the order from left to right of the text); among them, the new character is generated based on the generated character sequence, and after the new character is generated, the new character is added to the back of the generated character sequence.
[0089] For example, the position parameter form corresponding to the pseudo prompt text and the text of the user's consultation question in the embedding layer of the GPT-2 model is 《pseudo prompt text|text of the user's consultation question》, and the pseudo prompt text and the text of the user's consultation question (for example, X=t 1 ,t 2 ,…,t n) After the input GPT-2 model is encoded by the embedding layer, it outputs prompt encoding information and text encoding information e(t 1 ), e(t 2 ), …, e(t n ). The prompt encoding information and text encoding information are input into the GPT-2 model for text generation. Among them, the generated character sequence is S = X 1 , X 2 , …, X m-1 , and the character sequence to be generated is T = X m , X m+1 , …, X n . Then the generation probability of the m-th character is p(T|S) = ∏ n p(X m |X 1 , X 2 , …, X m-1 ). After the m-th character is generated, the m-th character is added to the sequence S. At this time, the sequence S = X 1 , X 2 , …, X m-1 , X m , and the sequence T = X m+1 , X m+2 , …, X n .
[0090] After generating the predicted text, the embedding layer parameters corresponding to the pseudo-prompt text in the embedding layer of the GPT-2 model are optimized through the loss function (i.e., training the weights of the masked embedding layer). When the similarity between the predicted text generated by the GPT-2 model and the corresponding user consultation question text is greater than the second preset threshold, the optimization of the embedding layer parameters corresponding to the pseudo-prompt text is stopped, and a text generation model is obtained.
[0091] Step 103: Compare the similarity between the text in the common consultation question candidate set and the question text in the intelligent knowledge base to obtain multiple similarities;
[0092] The predicted text generated in step 102 constitutes the common consultation question candidate set. In the embodiment of the present application, the beam search algorithm is used to select a preset number (for example, top-n) of common consultation questions from the common consultation question candidate set to implement the first screening of the common consultation question candidate set, delete the uncommon user consultation questions, and then compare the similarity between the text of the top-n common consultation questions and the text of the questions in the intelligent knowledge base to obtain n similarities.
[0093] Among them, the specific implementation process of comparing the similarity between the text of the preset number of common consultation questions and the text of the questions in the intelligent knowledge base is as follows:
[0094] Convert the text of a preset number of common consultation questions and the text of questions in the intelligent knowledge base into word vector sequences with a fixed length respectively. Through a deep learning algorithm, obtain the first semantic feature vector sequence and the second semantic feature vector sequence corresponding to the first word vector sequence and the second word vector sequence respectively. Among them, the first word vector sequence is the word vector sequence corresponding to the first text in the preset number of common consultation questions, the second word vector sequence is the word vector sequence corresponding to the second text in the intelligent knowledge base, and the lengths of the first word vector sequence and the second word vector sequence are the same; Determine the first global feature vector sequence and the second global feature vector sequence corresponding to the first semantic feature vector sequence and the second semantic feature vector sequence according to the attention mechanism, calculate the similarity between each global feature vector in the first global feature vector sequence and each global feature vector in the second global feature vector sequence to obtain a similarity matrix, and determine the similarity between each text in the first text and each text in the second text according to this similarity matrix.
[0095] For example, convert a preset number of common consultation question texts Q = {Q 1 , Q 2 , … Q i} and the text of questions q = {q 1 , q 2 , … q j} in the intelligent knowledge base into word vector sequences with a fixed length respectively using the BERT pre-training model, and input the word vector sequences of Q and q with the same length into two bidirectional long short-term memory network Bi-LSTM units respectively to extract the context semantic feature vector sequences h Q and h q . According to the attention mechanism, select the feature vectors of the last moment states of the semantic feature vector sequences of Q and q for vector splicing as background information. Among them, this background information contains the semantic information of all previous time states of Q and q. Then, through the fully connected layer, reduce the dimension to half, and calculate the global feature vector sequences S Q and S q with global features according to the attention mechanism. Use the text similarity formula to calculate the similarity matrix sim of the global feature vector sequences S Q and S q . Let S Qa and S qb be the a-th and b-th global feature vectors of S Q and S q respectively. Then the similarity of the a-th and b-th global feature vectors is sim Qaqb = S Qa · S qb .
[0096] Step 104: Add the text in the common consultation question candidate set with a similarity higher than the first preset threshold to the intelligent knowledge base.
[0097] In the embodiment of the present application, k-MAX Pooling is used to select K feature information points with the largest similarities from the similarity matrix sim to form a feature vector representing the matching of the question text. The feature vector is input into the fully connected layer, and finally, a softmax classifier is used to perform binary classification prediction of text semantic matching. The text with a prediction value m greater than the first preset threshold M is added to the intelligent knowledge base, and the text with a prediction value m less than the first preset threshold M is deleted and not selected for warehousing. In a possible implementation manner, the text with a prediction value greater than the first preset threshold can also be manually approved for warehousing.
[0098] In the specific implementation process, it is possible to effectively and truly predict the online user consultation behavior and generate common consultation question texts. The generated common consultation question texts are matched with the texts of the questions in the intelligent knowledge base to expand the question texts in the intelligent knowledge base, which can effectively improve the coverage rate of user consultation questions for the FAQ knowledge points in the intelligent knowledge base, thereby improving the accuracy of the intelligent robot's question answering and the service quality of the intelligent robot.
[0099] Based on the same inventive concept, the embodiment of the present application provides an intelligent knowledge base expansion device, which can implement the functions corresponding to the foregoing intelligent knowledge base expansion method. The intelligent knowledge base expansion device can be a hardware structure, a software module, or a combination of a hardware structure and a software module. The intelligent knowledge base expansion device can be implemented by a chip system, and the chip system can be composed of chips or can include chips and other discrete devices. Please refer to Figure 2 As shown, the intelligent knowledge base expansion device includes a processing module 201 and a model training module 202. Among them:
[0100] The processing module 201 is used to obtain the prefix of the user consultation text; wherein, the prefix of the user consultation text is the first character of the user consultation text or a combination of consecutive multiple characters starting from the first character;
[0101] The processing module 201 is further used to input the prefix of the user consultation text into the trained text generation model for text generation to obtain a common consultation question candidate set; wherein, the text generation model is obtained by training the GPT-2 model. The embedding layer parameters of the GPT-2 model include the embedding layer parameters corresponding to the pseudo-prompt text and the embedding layer parameters corresponding to the user consultation question text. The pseudo-prompt text is text without an actual task and is used to obtain prompt information during the text generation process. During the training of the GPT-2 model, the embedding layer parameters corresponding to the user consultation question text are masked;
[0102] The processing module 201 is further configured to compare the text in the common consultation question candidate set with the text of the questions in the intelligent knowledge base to obtain multiple similarity degrees.
[0103] The processing module 201 is further configured to add the text with a similarity degree higher than a first preset threshold in the common consultation question candidate set to the intelligent knowledge base.
[0104] In a possible implementation manner, the intelligent knowledge base expansion device further includes a model training module 202, configured to:
[0105] Obtain the user consultation question text in a preset domain.
[0106] Input the pseudo prompt text and the user consultation question text into the embedding layer of the GPT-2 model for encoding to obtain the prompt encoding information corresponding to the pseudo prompt text and the text encoding information corresponding to the user consultation question text, so that the GPT-2 model generates a prediction text based on the prompt encoding information and the text encoding information; wherein, the prompt encoding information is used to prompt the characters to be generated.
[0107] Based on the prediction text, optimize the embedding layer parameters corresponding to the pseudo prompt text in the embedding layer of the GPT-2 model through a loss function until the similarity between the prediction text generated by the GPT-2 model and the user consultation question text is greater than a second preset threshold, and then stop optimizing the embedding layer parameters corresponding to the pseudo prompt text to obtain the text generation model.
[0108] In a possible implementation manner, the model training module 202 is specifically configured to:
[0109] Generate each character in sequence according to the standard grammar order based on the prompt encoding information and the text encoding information; wherein, the new character is generated based on the generated character sequence, and after the new character is generated, the new character is added to the back of the generated character sequence.
[0110] In a possible implementation manner, the processing module 201 is specifically configured to:
[0111] Preprocess the user consultation question text to obtain a target text; wherein, the preprocessing includes deleting the text containing sensitive vocabulary, incomplete text, picture URLs, special characters, converting the complex writing to the simplified writing, and converting the English capital letters to the lowercase letters, the sensitive vocabulary is the vocabulary included in a preset sensitive database, and the incomplete text is the text with a text grammar different from the standard grammar.
[0112] Input the pseudo prompt text and the target text into the embedding layer of the GPT-2 model for encoding.
[0113] In a possible implementation, the processing module 201 is specifically configured to:
[0114] Crawl user consultation text data;
[0115] Count the text prefixes in the user consultation text data;
[0116] Determine the number of occurrences of each text prefix;
[0117] Sort the text prefixes according to the number of occurrences to obtain a priority order; among them, the text prefix with the most occurrences has the highest priority;
[0118] Obtain the user consultation text prefix for text generation according to the priority order.
[0119] In a possible implementation, the processing module 201 is specifically configured to:
[0120] Select a preset number of common consultation questions from the candidate set of frequently asked questions through a beam search algorithm;
[0121] Compare the texts of the preset number of common consultation questions with the texts of the questions in the intelligent knowledge base for similarity.
[0122] In a possible implementation, the processing module 201 is specifically configured to:
[0123] Convert the texts of the preset number of common consultation questions and the texts of the questions in the intelligent knowledge base into fixed-length word vector sequences respectively;
[0124] Obtain a first semantic feature vector sequence and a second semantic feature vector sequence corresponding to the first word vector sequence and the second word vector sequence respectively through a deep learning algorithm; wherein, the first word vector sequence is the word vector sequence corresponding to the first text in the preset number of common consultation questions, the second word vector sequence is the word vector sequence corresponding to the second text in the intelligent knowledge base, and the lengths of the first word vector sequence and the second word vector sequence are the same;
[0125] Determine a first global feature vector sequence and a second global feature vector sequence corresponding to the first semantic feature vector and the second semantic feature vector sequence according to the attention mechanism;
[0126] Calculate the similarity between each global feature vector in the first global feature vector sequence and each global feature vector in the second global feature vector to obtain a similarity matrix;
[0127] Determine the similarity between each text in the first text and each text in the second text according to the similarity matrix.
[0128] All relevant content of each step involved in the embodiments of the foregoing intelligent knowledge base expansion method can be cited in the function description of the corresponding functional modules of the intelligent knowledge base expansion device in the embodiments of the present application, and will not be elaborated herein.
[0129] The division of modules in the embodiments of the present application is illustrative, and is only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present application, the various functional modules may be integrated in one processor, or may exist separately physically, or two or more modules may be integrated in one module. The above integrated modules may be implemented in the form of hardware or in the form of software functional modules.
[0130] Based on the same inventive concept, an embodiment of the present application provides an electronic device. Please refer to Figure 3 As shown, the electronic device includes at least one processor 301 and a memory 302 connected to the at least one processor. In the embodiments of the present application, the specific connection medium between the processor 301 and the memory 302 is not limited. Figure 3 Taking the connection between the processor 301 and the memory 302 through the bus 300 as an example, the bus 300 is Figure 3 shown in thick lines. The connection methods between other components are only for illustrative purposes and are not limited thereto. The bus 300 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 3 only one thick line is shown herein, but it does not mean that there is only one bus or one type of bus.
[0131] In the embodiments of the present application, the memory 302 stores instructions executable by the at least one processor 301. By executing the instructions stored in the memory 302, the at least one processor 301 can execute the steps included in the foregoing intelligent knowledge base expansion method.
[0132] Among them, the processor 301 is the control center of the electronic device. It can connect various parts of the entire electronic device through various interfaces and circuits. By running or executing instructions stored in the memory 302 and calling data stored in the memory 302, it can perform various functions of the electronic device and process data, thereby monitoring the electronic device as a whole. Optionally, the processor 301 may include one or more processing units. The processor 301 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 301 either. In some embodiments, the processor 301 and the memory 302 may be implemented on the same chip. In some embodiments, they may also be separately implemented on independent chips.
[0133] The processor 301 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the intelligent knowledge base expansion method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0134] As a non-volatile computer-readable storage medium, the memory 302 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 302 may include at least one type of storage medium. For example, it may include flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (RAM), a static random access memory (SRAM), a programmable read-only memory (PROM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic memory, a magnetic disk, an optical disk, and so on. The memory 302 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 302 in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.
[0135] By programming the design of the processor 301, the code corresponding to the intelligent knowledge base expansion method introduced in the foregoing embodiments can be solidified into the chip, so that the chip can execute the steps of the intelligent knowledge base expansion method when running. How to program the design of the processor 301 is a well-known technology to those skilled in the art and will not be elaborated here.
[0136] Based on the same inventive concept, an embodiment of the present application also provides a computer-readable storage medium, which stores computer instructions. When the computer instructions run on a computer, the computer is caused to execute the steps of the intelligent knowledge base expansion method as described above.
[0137] In some possible implementation manners, each aspect of the intelligent knowledge base expansion method provided by the present application can also be implemented in the form of a program product, which includes program code. When the program product runs on an electronic device, the program code is used to cause the detection device to execute the steps in the intelligent knowledge base expansion method according to various exemplary embodiments of the present application described above in this specification.
[0138] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0139] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0140] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions in the process Figure 1 one process or multiple processes and / or blocksFigure 1 The functions specified in one or more boxes.
[0141] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in one Figure 1 one process or more processes and / or boxes Figure 1 step of the functions specified in one or more boxes. Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.
Claims
1. A method for expanding an intelligent knowledge base, characterized in that: The method comprises: Obtaining a user consultation text prefix; wherein the user consultation text prefix is the first word in the user consultation text or a combination of multiple consecutive words starting from the first word; Input the user consultation text prefix into a trained text generation model to generate text and obtain a candidate set of common consultation questions; wherein the text generation model is obtained by training a GPT-2 model, and the embedding layer parameters of the GPT-2 model include embedding layer parameters corresponding to pseudo prompt text and embedding layer parameters corresponding to user consultation question text, wherein the pseudo prompt text is text without actual tasks and is used to obtain prompt information in the text generation process, and during the training of the GPT-2 model, the embedding layer parameters corresponding to the user consultation question text are shielded; Comparing the texts in the candidate set of common consulting questions with the texts of questions in the intelligent knowledge base for similarity, and obtaining multiple similarities; The texts in the candidate set of common consulting questions with a similarity higher than a first preset threshold are added to the intelligent knowledge base.
2. The method according to claim 1, characterized in that The text generation model is obtained in the following way: Get the user consultation question text in the preset field; Inputting the pseudo prompt text and the user consultation question text into the embedding layer of the GPT-2 model for encoding, obtaining prompt encoding information corresponding to the pseudo prompt text and text encoding information corresponding to the user consultation question text, so that the GPT-2 model generates a predicted text based on the prompt encoding information and the text encoding information; wherein the prompt encoding information is used to prompt the word to be generated; Based on the predicted text, the embedding layer parameters corresponding to the pseudo prompt text in the embedding layer of the GPT-2 model are optimized through a loss function until the similarity between the predicted text generated by the GPT-2 model and the user consultation question text is greater than a second preset threshold, and then the embedding layer parameters corresponding to the pseudo prompt text are stopped to obtain the text generation model.
3. The method according to claim 2, characterized in that The GPT-2 model generates predicted text based on the prompt encoding information and the text encoding information, including: Generate each word in sequence according to the standard grammatical order based on the prompt encoding information and the text encoding information; wherein, the new word is generated based on the generated word sequence, and after the new word is generated, the new word is added to the end of the generated word sequence.
4. The method according to claim 2, characterized in that The step of inputting the pseudo prompt text and the user consultation question text into the embedding layer of the GPT-2 model for encoding includes: Preprocessing the user consultation question text to obtain a target text; wherein the preprocessing includes deleting text containing sensitive words, incomplete text, image URLs, special characters, converting traditional Chinese characters into abbreviations, and converting English uppercase characters into lowercase characters, wherein the sensitive words are words contained in a preset sensitive database, and the incomplete text is text whose text syntax is different from the standard syntax; The pseudo prompt text and the target text are input into the embedding layer of the GPT-2 model for encoding.
5. The method according to claim 1, characterized in that The obtaining of the user consultation text prefix includes: Crawl user consultation text data; Collecting statistics on text prefixes in the user consultation text data; Determine the number of times each text prefix occurs; The text prefixes are prioritized according to the number of occurrences; the text prefix with the most occurrences is the text prefix with the highest priority; Get the user consultation text prefixes for text generation according to the priority order.
6. The method according to claim 1, characterized in that The similarity comparison between the text in the candidate set of common consulting questions and the text of questions in the intelligent knowledge base includes: Selecting a preset number of common consulting questions from the candidate set of common consulting questions by using a beam search algorithm; The texts of the preset number of common consulting questions are compared for similarity with the texts of the questions in the intelligent knowledge base.
7. The method according to claim 6, characterized in that The similarity comparison of the preset number of common consulting question texts with the texts of questions in the intelligent knowledge base includes: Converting the preset number of common consulting question texts and the texts of the questions in the intelligent knowledge base into word vector sequences with fixed lengths respectively; Acquire a first semantic feature vector sequence and a second semantic feature vector sequence corresponding to the first word vector sequence and the second word vector sequence respectively by a deep learning algorithm; wherein the first word vector sequence is a word vector sequence corresponding to a first text in the preset number of common consulting question texts, and the second word vector sequence is a word vector sequence corresponding to a second text in the intelligent knowledge base, and the first word vector sequence and the second word vector sequence have the same length; Determine, according to the attention mechanism, a first global feature vector sequence and a second global feature vector sequence corresponding to the first semantic feature vector and the second semantic feature vector sequence; Calculating the similarity between each global feature vector in the first global feature vector sequence and each global feature vector in the second global feature vector sequence to obtain a similarity matrix; The similarity between each text in the first text and each text in the second text is determined according to the similarity matrix.
8. An intelligent knowledge base expansion device, characterized in that: The device comprises: A processing module, used for obtaining a user consultation text prefix; wherein the user consultation text prefix is the first word in the user consultation text or a combination of multiple consecutive words starting from the first word; The processing module is further used to input the user consultation text prefix into a trained text generation model to generate text and obtain a candidate set of common consultation questions; wherein the text generation model is obtained by training a GPT-2 model, and the embedding layer parameters of the GPT-2 model include embedding layer parameters corresponding to pseudo prompt text and embedding layer parameters corresponding to user consultation question text, wherein the pseudo prompt text is text without actual tasks and is used to obtain prompt information in the text generation process, and during the training of the GPT-2 model, the embedding layer parameters corresponding to the user consultation question text are shielded; The processing module is further used to compare the texts in the candidate set of common consulting questions with the texts of questions in the intelligent knowledge base to obtain multiple similarities; The processing module is further configured to add texts in the common consulting question candidate set whose similarity is higher than a first preset threshold to the intelligent knowledge base.
9. An electronic device, characterized in that: include: A memory for storing program instructions; A processor is used to call the program instructions stored in the memory, and execute the steps included in the method according to any one of claims 1 to 7 according to the obtained program instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Question and answer processing method and device, equipment and medium
CN111382255A
Question retrieval method and device applied to intelligent question-answering system, equipment and medium
CN113590789A