Question generation method, device, computer equipment and storage medium

By matching and vector transformation from the dialogue database to generate question sentences, the problem of insufficient flexibility and diversity in question generation in the intelligent dialogue system is solved, and a high accuracy and fluency question generation is achieved.

CN111666392BActive Publication Date: 2025-08-22CHINA PING AN LIFE INSURANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010350584.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-04-28
Publication Date
2025-08-22
Estimated Expiration
2040-04-28

AI Technical Summary

Technical Problem

In the existing intelligent dialogue system, the method of question generation is poor in flexibility and diversity, and it is easy to generate meaningless and low-information dialogue sentence patterns, affecting the fluency of dialogue.

Method used

By matching the input text information from the dialogue database, obtaining the above dialogue information and replacing it, using vector transformation and bidirectional encoder processing, the following dialogue question is generated closely related to the input text information.

Benefits of technology

It improves the accuracy and fluency of question generation, ensuring that the generated question is highly matched with the input text information and is practical.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111666392B_ABST
    Figure CN111666392B_ABST
Patent Text Reader

Abstract

The present invention discloses a question generation method, apparatus, computer device, and storage medium. The method comprises: obtaining input text information; matching the input text information from a conversation database to obtain preceding conversation information matching the input text information and following conversation information corresponding to the preceding conversation information; replacing the preceding conversation information with the input text information to obtain input preceding conversation information; performing vector conversion on the preceding and following conversation information to obtain an input preceding vector and an following conversation vector; and processing the preceding and following conversation vectors using a bidirectional encoder to obtain a following conversation question corresponding to the input text information. The technical solution of the present invention solves the problem of poor fluency in generated questions by enhancing the close connection between the input text information and the following conversation question, thereby ensuring that the generated questions are fluent and practical.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a question generation method, apparatus, computer equipment, and storage medium. Background Art

[0002] Nowadays, intelligent dialogue systems are widely used in our daily lives. Questions play a very important role in intelligent dialogue systems. Whether the questions generated by the intelligent dialogue system can stimulate the interest of the person being spoken to and make the conversation proceed smoothly determines the practicality of the intelligent dialogue system.

[0003] Traditional question generation methods are mainly template-based. In intelligent dialogue systems, questions are generated by matching relevant topics using pre-defined templates. This method requires a lot of manpower to implement, has poor flexibility and diversity, and is prone to generating meaningless and low-information dialogue sentences, such as "Are you okay? Really? Are you there?" Summary of the Invention

[0004] Embodiments of the present invention provide a question generation method, apparatus, computer device, and storage medium to solve the problem of poor fluency in generated question sentences.

[0005] A question generation method, comprising:

[0006] Get input text information;

[0007] Matching the input text information from a conversation database to obtain previous conversation information matching the input text information and following conversation information corresponding to the previous conversation information;

[0008] Replacing the previous dialogue information with the input text information to obtain input previous dialogue information, and performing vector conversion on the input previous dialogue information and the next dialogue information to obtain an input previous dialogue vector and a next dialogue vector;

[0009] The input context vector and the context dialogue vector are processed by a bidirectional encoder to obtain a context dialogue question corresponding to the input text information.

[0010] A question generation device, comprising:

[0011] Information input module, used to obtain input text information;

[0012] An information matching module is used to match the input text information from a conversation database to obtain previous conversation information that matches the input text information and next conversation information corresponding to the previous conversation information;

[0013] a vector conversion module, configured to replace the preceding dialogue information with the input text information to obtain preceding input information, and to perform vector conversion on the preceding input information and the following dialogue information to obtain preceding input vectors and following dialogue vectors;

[0014] The vector processing module is used to process the input context vector and the context dialogue vector through a bidirectional encoder to obtain a context dialogue question corresponding to the input text information.

[0015] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned question generation method are implemented.

[0016] A computer-readable storage medium stores a computer program, which implements the steps of the question generation method when executed by a processor.

[0017] In the above-mentioned question generation method, apparatus, computer device, and storage medium, input text information is matched from a dialogue database to obtain previous dialogue information matching the input text information and subsequent dialogue information corresponding to the previous dialogue information. The subsequent dialogue information that best matches the input text information can be obtained from the dialogue database, so that the finally generated question sentence has a high degree of match with the input text information, thereby improving the accuracy of the question generation method. The previous dialogue information is replaced based on the input text information to obtain input previous information, and vector conversion is performed on the input previous information and the subsequent dialogue information to obtain an input previous vector and a subsequent dialogue vector. The input text information and the subsequent dialogue information can be converted into an input previous vector and a subsequent dialogue vector through vector conversion, thereby obtaining corresponding vector features from the input previous vector and the subsequent dialogue vector, thereby improving the fluency of the question sentence based on the vector features. The input previous vector and the subsequent dialogue vector are processed by a bidirectional encoder to obtain a subsequent dialogue question sentence corresponding to the input text information, thereby improving the close connection between the input text information and the subsequent dialogue question sentence, thereby making the generated question sentence practical. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0019] Figure 11 is a schematic diagram of an application environment of a question generation method according to an embodiment of the present invention;

[0020] Figure 2 is a flow chart of a question generation method according to an embodiment of the present invention;

[0021] Figure 3 is a flow chart of step S20 in the question generation method according to one embodiment of the present invention;

[0022] Figure 4 is a flow chart of a question generation method according to an embodiment of the present invention;

[0023] Figure 5 is a flow chart of a question generation method according to an embodiment of the present invention;

[0024] Figure 6 is a flow chart of a question generation method according to an embodiment of the present invention;

[0025] Figure 7 is a schematic diagram of a question generation device according to an embodiment of the present invention;

[0026] Figure 8 FIG. 1 is a schematic diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0028] The question generation method provided in this application can be applied to Figure 1In the application environment shown, the application environment is specifically an intelligent dialogue system, which includes a server and a client, wherein the server and the client are connected through a network, which can be a wired network or a wireless network. The client specifically includes but is not limited to various personal computers, laptops, smart phones and tablets and portable wearable devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers. The server obtains input text information; matches the input text information from a dialogue database to obtain previous dialogue information matching the input text information and subsequent dialogue information corresponding to the previous dialogue information, and can obtain the subsequent dialogue information that best matches the input text information from the dialogue database, so that the finally generated question sentence can have a high degree of matching with the input text information, thereby improving the accuracy of the question generation method; according to the input text information, the previous dialogue information is replaced to obtain input previous information, and vector conversion is performed on the input previous information and the subsequent dialogue information respectively to obtain an input previous vector and a subsequent dialogue vector, and can convert the input text information and the subsequent dialogue information into an input previous vector and a subsequent dialogue vector by vector conversion, so that corresponding vector features can be obtained from the input previous vector and the subsequent dialogue vector, thereby improving the fluency of the question sentence based on the vector features; through a bidirectional encoder, the input previous vector and the subsequent dialogue vector are processed to obtain a subsequent dialogue question sentence corresponding to the input text information, thereby improving the close connection between the input text information and the subsequent dialogue question sentence, so that the generated question sentence has practicality.

[0029] In one embodiment, if Figure 2 As shown, a question generation method is provided, which is applied in Figure 1 The server in the example is used as an example, specifically including steps S10 to S40, which are detailed as follows:

[0030] S10: Acquire input text information.

[0031] The input text information refers to the text information input by the user, that is, the server obtains the text information input by the user through the client.

[0032] As an example, the input text information may be text information directly input by a user through a client and obtained by the server. The acquisition of the input text information includes but is not limited to acquisition through a keyboard or a touch screen display.

[0033] As an example, the input text information can be text information obtained by the server obtaining the voice information input by the user through the client and converting the input voice information into text. The voice information can be obtained through a microphone, which can improve the user experience. It should be noted that after the server obtains the voice information, it automatically converts the voice information into corresponding text information. For example, the voice information obtained is "I like to eat hamburgers", and the voice information is automatically converted into the text information "I like to eat hamburgers".

[0034] S20: Match the input text information from the dialogue database to obtain previous dialogue information that matches the input text information and following dialogue information that corresponds to the previous dialogue information.

[0035] The conversation database is a pre-created database used to store raw conversation data. This raw conversation data is retrieved from a user-defined database. The conversation data is filtered to obtain a conversation database containing question sentences. The conversation database containing question sentences is a database created by deriving question sentence conversation data from the user-defined conversation database. Previous conversation information refers to information about the previous conversation in the conversation database. Next conversation information refers to information about the next conversation in the conversation database that corresponds to the input text and is the next conversation in the question sentence.

[0036] As an example, after receiving input text, the server matches the input text against the conversation database to obtain the preceding conversation information that matches the input text. In this example, the server matches the input text using a full-text search engine. Optionally, the full-text search engine can be a search engine such as Lucene, Solr, or Elastic Search.

[0037] Preferably, the full-text search engine can be a Lucene engine. It should be noted that the server can use the built-in similarity matching algorithm in the Lucene engine to perform similarity calculations on the input text information in the conversation database to obtain previous conversation information that matches the input text information. Optionally, the similarity matching algorithm can be a vector space model algorithm or a BM25 algorithm.

[0038] As an example, the server uses the Lucene engine's similarity matching algorithm to calculate the word frequency and inverse word frequency of each word in the input text in the conversation database. Using these frequencies and inverse word frequencies, the server then obtains the conversation data with the highest similarity to the input text. Furthermore, the server obtains the preceding conversation information and the corresponding following conversation information from the conversation data with the highest similarity to the input text.

[0039] For example, the server uses the BM25 algorithm in the Lucene engine to calculate the word frequency and inverse text frequency of each word in the input text information "I like to eat hamburgers" in the conversation database, and matches the conversation information with the highest similarity to "I like to eat steamed buns." In the conversation database, the corresponding conversation information is "Which steamed bun do you think is delicious?"

[0040] S30: replacing the previous conversation information with the input text information to obtain input previous context information, and performing vector conversion on the input previous context information and the next conversation information to obtain an input previous context vector and a next conversation vector.

[0041] The input context information refers to the information obtained by replacing the input text information with the input context information. The input context vector refers to the vector information obtained by converting the input context information into a vector. The following conversation vector refers to the vector information obtained by converting the following conversation information into a vector. Optionally, the vector conversion can be performed using a word2vec model, a glove model, an ELMo model, or a BERT model. For example, the vector conversion can be performed using a word2vec model or a BERT model, which can further describe word-level, sentence-level, and conversation-level relationship features.

[0042] Specifically, the server uses a full-text search engine to match the input text information, obtaining the preceding conversation information that matches the input text information, and the following conversation information that corresponds to the preceding conversation information. First, the preceding conversation information is replaced with the input text information to obtain the preceding input information. Then, the preceding input information and the following conversation information are converted to word vectors using the word2vec model to obtain the preceding word vector and the following word vector. Next, the transformer in the BERT model is used as a feature extractor to extract features from the word vectors in the preceding and following word vectors. The preceding and following word vectors are then converted to sentence vectors using the BERT model to obtain the preceding input vector and the following conversation vector. Word vector conversion refers to the process of extracting words or phrases from the text content of the preceding input information and the following conversation information, and then extracting feature vectors for these words or phrases. A feature vector refers to the mapping between words or phrases in the text content of the preceding input information and the words or phrases in the following conversation information. For example, the feature vector of the word or phrase "I like to eat hamburgers" in the input information above is "I like to eat hamburgers", and the feature vector of the word or phrase "Which shaomai do you think is good?" in the following dialogue information is "Which shaomai do you think is good?"

[0043] For example, based on the input text "I like to eat hamburgers," the preceding conversation information "I like to eat shaomai" matched from the conversation database is replaced to obtain the input preceding text "I like to eat hamburgers." Furthermore, the word2vec model is used to convert the text content of the preceding input information "I like to eat hamburgers" and the text content of the following conversation information "Which shaomai do you think is good to eat?" into word vectors. The transformer in the BERT model is used as a feature extractor to extract the word vectors from "I like to eat hamburgers" and "Which shaomai do you think is good to eat?" The BERT model then performs sentence vector conversion on the preceding and following word vectors, obtaining the input preceding vector "I like to eat hamburgers" and the following conversation vector "Which shaomai do you think is good to eat?"

[0044] S40: Process the input context vector and the context vector through a bidirectional encoder to obtain a context question corresponding to the input text information.

[0045] The bidirectional encoder refers to an encoder capable of bidirectionally fusion encoding sentence vectors. Optionally, the bidirectional encoder can be a bidirectional recurrent neural network model (Gated Recurrent Unit, GRU). The following dialogue question refers to the following dialogue data corresponding to the input text information.

[0046] As an example, the server performs vector conversion on the input context and the following conversation information to obtain the context vector "I like hamburgers" and the conversation vector "Which shaomai do you think is delicious?". First, the server uses a bidirectional encoder to encode the context vector "I like hamburgers" and the conversation vector "Which shaomai do you think is delicious?" to obtain the context vector "I like hamburgers" and the conversation vector "Which shaomai do you think is delicious?". Then, the server calculates the vector dimensions of the context vector "I like hamburgers" and the conversation vector "Which shaomai do you think is delicious?" to obtain the conversation sentence vector "I like hamburgers: Which shaomai do you think is delicious?". Next, the bidirectional encoder decodes the conversation sentence vector "I like hamburgers: Which shaomai do you think is delicious?" to obtain the conversation question "Which hamburgers do you think is delicious?" corresponding to the input text "I like hamburgers." Calculating the vector dimensions refers to adding the vector dimensions of the context vector to the vector dimensions of the conversation vector. The following dialogue question sentence corresponding to the input text information refers to a question sentence generated based on the input text information.

[0047] In this embodiment, input text information is matched against a conversation database to obtain preceding conversation information that matches the input text information and following conversation information corresponding to the preceding conversation information. The following conversation information that best matches the input text information can be obtained from the conversation database, ensuring that the generated question has a high degree of match with the input text information, thereby improving the accuracy of the question generation method. The preceding conversation information is replaced with the input text information to obtain the preceding information. The preceding information and the following conversation information are then vector-converted to obtain an input preceding vector and a following conversation vector. The input text information and the following conversation information are converted into an input preceding vector and a following conversation vector through vector conversion, thereby obtaining corresponding vector features from the input preceding vector and the following conversation vector. The fluency of the question is thereby improved based on the vector features. The preceding vector and the following conversation vector are processed by a bidirectional encoder to obtain a following conversation question corresponding to the input text information. This improves the close connection between the input text information and the following conversation question, thereby enhancing the fluency of the generated question.

[0048] In one embodiment, if Figure 3 As shown, in step S20, the input text information is matched from the conversation database to obtain the previous conversation information matching the input text information and the following conversation information corresponding to the previous conversation information, including:

[0049] S21: Calculate the similarity between the input text information and each original conversation data in the conversation database through the full-text search engine to obtain the similarity matching value corresponding to each original conversation data.

[0050] The full-text search engine may be a Lucene engine, a Solr engine, an Elasticsearch engine, etc. For example, the full-text search engine may be a Lucene engine, which provides a complete query engine and index engine.

[0051] Specifically, the server uses a full-text search engine to calculate similarity between each conversation data item containing a question and the input text information in the conversation database, obtaining a similarity match value between the input text information and each conversation data item. The similarity calculation for the input text information is performed using the built-in similarity algorithm in the full-text search engine.

[0052] S22: Extract N original conversation data with similarity matching values ​​exceeding a similarity threshold from the conversation database, and obtain N target original data, where N is a positive integer.

[0053] N is a user-defined number, a positive integer. The similarity threshold is a user-defined parameter related to similarity. Specifically, the server uses a full-text search engine to perform a similarity match between the input text and all conversation data in the conversation database. The server then extracts conversation data with similarity match values ​​greater than the similarity threshold.

[0054] S23: Determine, based on a preset probability, the preceding dialogue information that matches the input text information and the following dialogue information that corresponds to the preceding dialogue information from the N target original data.

[0055] The preset probability is a random value greater than 0 and less than 1. Specifically, according to the preset probability, N different random values ​​of the preset probability are generated corresponding to the N conversation data. It can be understood that the values ​​of the N preset probabilities are random values ​​greater than 0 and less than 1.

[0056] As an example, based on N preset probabilities, previous conversation information matching the input text information is selected from the N extracted conversation data, and following conversation information corresponding to the previous conversation information is determined, so that the generated question sentences have diversity.

[0057] In this embodiment, the server uses a full-text search engine to calculate the similarity between the input text information and each conversation data in the conversation database to improve the accuracy of the final generated subsequent conversation questions. N conversation data whose similarity with the input text information exceeds a similarity threshold are extracted from the conversation database. The conversation data are extracted based on the similarity threshold to further improve the accuracy of the final generated subsequent conversation questions. Based on preset probabilities, the server determines the previous conversation information that matches the input text information and the subsequent conversation information corresponding to the previous conversation information from the N conversation data, thereby improving the diversity of question generation.

[0058] In one embodiment, if Figure 4 As shown, before step S30, the previous dialogue information is replaced according to the input text information to obtain the input previous information, and the input previous information and the following dialogue information are respectively vector-converted to obtain the input previous vector and the following dialogue vector, including:

[0059] S301: Length identification is performed on the input preceding information and the following dialogue information to obtain a dialogue length result, wherein the dialogue length result is byte length information of the input preceding information and the following dialogue information.

[0060] The conversation length result refers to the byte length result of the input preceding information and the following conversation information. Specifically, the server performs byte recognition on the input preceding information and the following conversation information, and inputs the conversation length results of the preceding information and the following conversation information respectively.

[0061] As an example, the server performs byte recognition on the input information above "I like to eat hamburgers" and the following dialogue information "Which shaomai do you think is delicious?", and obtains byte length results A bytes and B bytes, and determines that the dialogue length result corresponding to the input information above is A, and the dialogue length result corresponding to the following dialogue information is B.

[0062] S302: If the resulting conversation length is less than the preset conversation length, the lengths of the input preceding information and the following conversation information are padded according to the preset conversation length.

[0063] The preset conversation length is the user-defined conversation data length. Specifically, the server compares the conversation length results for the preceding and following input messages with the preset conversation length. If any of the conversation length results for the preceding or following input messages are shorter than the preset conversation length, the corresponding conversation length of the preceding or following input message that is shorter than the preset conversation length is padded with zeros to ensure that the conversation length of the preceding or following input message equals the preset conversation length.

[0064] As an example, the preset conversation length is C, the input context conversation length is A, and the conversation length corresponding to the following conversation information is B. If A or B is less than C, the input context conversation length A or the following conversation length B is equal to the preset conversation length C by adding 0.

[0065] S303: If the result of the conversation length is greater than the preset conversation length, the lengths of the input preceding information and the following conversation information are cut off according to the preset conversation length and the preset cutting rule.

[0066] The preset interception rule refers to a user-defined rule for intercepting the length of a conversation. Optionally, the preset interception rule can intercept the length of the conversation from the beginning to the end of the preceding and following conversations, or from the end to the beginning of the preceding and following conversations. Preferably, the preset interception rule can intercept the length of the conversation from the end to the beginning of the preceding and following conversations. It should be noted that intercepting the length of the conversation from the end to the beginning of the preceding and following conversations, based on human language habits, can more likely capture the primary information expressed in a sentence.

[0067] In this embodiment, the length of the input preceding and following dialogue information is identified to obtain a dialogue length result. Furthermore, if the dialogue length result is less than a preset dialogue length, the lengths of the input preceding and following dialogue information are padded according to the preset dialogue length, thereby improving the efficiency of vector conversion when the input preceding and following dialogue information are converted to vectors. If the dialogue length result is greater than the preset dialogue length, the lengths of the input preceding and following dialogue information are truncated according to a preset truncation rule based on the preset dialogue length, thereby improving the efficiency of vector conversion when the input preceding and following dialogue information are converted to vectors, thereby improving the efficiency of question generation.

[0068] In one embodiment, if Figure 5 As shown, the question generation method further includes:

[0069] S50: updating the number of subsequent dialogue questions corresponding to the input text information, and determining whether the number of subsequent dialogue questions reaches a preset number of cycles M, where M is a positive integer.

[0070] The preset number of loops M refers to the number of step executions set by the user. Specifically, the server updates the number of subsequent dialogue questions corresponding to the input text information and, based on the preset number of loops M, determines whether the number of subsequent dialogue questions has reached the preset number of loops M. It is understood that the server obtains subsequent dialogue questions corresponding to the input text information based on the preset number of loops M to update the number of subsequent dialogue questions, and then determines the subsequent dialogue question that best matches the input text information from the updated number of corresponding subsequent dialogue questions, thereby further improving the accuracy of the subsequent dialogue questions.

[0071] S60: If the number of following dialogue questions does not reach the preset number of loops M, the process returns to executing the process of extracting N original dialogue data with similarity matching values ​​exceeding a similarity threshold from the dialogue database to obtain N target original data, where N is a positive integer; determining, based on a preset probability, the preceding dialogue information that matches the input text information and the following dialogue information corresponding to the preceding dialogue information from the N target original data; replacing the preceding dialogue information with the input text information to obtain the preceding input information, performing vector conversion on the preceding input information and the following dialogue information respectively to obtain the preceding input vector and the following dialogue vector; processing the preceding input vector and the following dialogue vector through a bidirectional encoder to obtain the following dialogue questions corresponding to the input text information, until M dialogue questions to be confirmed that match the input text information are obtained.

[0072] Specifically, the number of subsequent dialogue questions is determined based on a preset number of loops M. When it is determined that the number of subsequent dialogue questions does not reach the preset number of loops M, the process returns to executing, based on a preset probability, determining, from N dialogue data, previous dialogue information that matches the input text information and subsequent dialogue information corresponding to the previous dialogue information; replacing the previous dialogue information with the input text information to obtain input previous information, performing vector conversion on the input previous information and the subsequent dialogue information respectively to obtain an input previous vector and a subsequent dialogue vector; processing the input previous vector and the subsequent dialogue vector through a bidirectional encoder to obtain subsequent dialogue questions corresponding to the input text information, until M dialogue questions to be confirmed that match the input text information are obtained.

[0073] S70: Based on the input text information, similarity calculation is performed on the M dialogue questions to be confirmed using the Jaccard coefficient, and the following dialogue question corresponding to the input text information is updated.

[0074] The Jaccard coefficient calculates the similarity between the input text and the question to be confirmed using the word features of both. Specifically, the server calculates the ratio of the intersection and union of the words in the input text and M questions to be confirmed, generating M similarity results. The server then sorts these M similarity results to obtain a ranked similarity result. Based on the top-ranked similarity result, the server determines the next question corresponding to the input text.

[0075] As an example, the server obtains M similarity results by obtaining the ratio of the intersection and union of the input text information and the words in M ​​to-be-confirmed dialogue questions. For example, the server obtains the ratio of the intersection and union of the input text information "I like to eat hamburgers" and one of the M to-be-confirmed dialogue questions "Which shaomai do you think is delicious?" to obtain a similarity result, that is, the ratio of the intersection and union of the input text information and the words in the to-be-confirmed dialogue question. The larger the ratio of the intersection and union of the words, the higher the similarity between the input text information and the to-be-confirmed dialogue question.

[0076] In this embodiment, the server obtains M dialogue questions to be confirmed that match the input text information according to a preset number of cycles M, calculates the similarity of the M dialogue questions to be confirmed based on the input text information and the Jaccard coefficient, and determines the following dialogue question corresponding to the input text information, which can improve the accuracy of generating question sentences.

[0077] In one embodiment, if Figure 6As shown, before matching the input text information from the dialogue database to obtain the previous dialogue information matching the input text information and the following dialogue information corresponding to the previous dialogue information, the question generation method further includes:

[0078] S201: Obtain all conversation data to be filtered from the original database.

[0079] The original database is a user-defined database that stores all conversation data to be screened. This data includes statements or questions. Specifically, the database address of the original database is first obtained. The original database is found using the database address, and then all conversation data to be screened is retrieved from the original database.

[0080] S202: Using a full-text search engine, the conversation data to be filtered is filtered according to a preset filtering method to obtain original conversation data, and a conversation database is created based on the original conversation data.

[0081] The full-text search engine may be Lucene. The preset filtering method refers to filtering out conversation data in the conversation data to be filtered where the following conversation is a question. The original conversation data refers to the conversation data where the following conversation is a question. The conversation database refers to the database of stored conversation data where the following conversation is a question. The full-text search engine refers to a full-text search engine that searches all acquired conversation data to be filtered and filters the conversation data according to the preset filtering method.

[0082] In this embodiment, the server obtains all the dialogue data to be filtered from the original database, further narrowing the scope of the dialogue data and improving the efficiency of generating question sentences. Through the full-text search engine, the dialogue data to be filtered is filtered according to a preset filtering method to obtain the original dialogue data, and a dialogue database is created based on the original dialogue data. By pre-creating a new dialogue database with the following dialogue as question sentences, the efficiency of generating question sentences is improved.

[0083] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0084] In one embodiment, a question generation device is provided, which corresponds to the question generation method in the above embodiment. Figure 7 As shown, the question generation device includes: an information input module 10, an information matching module 20, a vector conversion module 30, and a vector processing module 40. The functional modules are described in detail as follows:

[0085] Information input module 10, used to obtain input text information;

[0086] An information matching module 20 is used to match the input text information from the conversation database to obtain the previous conversation information that matches the input text information and the following conversation information that corresponds to the previous conversation information;

[0087] A vector conversion module 30 is configured to replace the previous dialogue information with the input text information to obtain the input previous dialogue information, and to perform vector conversion on the input previous dialogue information and the next dialogue information to obtain the input previous dialogue vector and the next dialogue vector;

[0088] The vector processing module 40 is used to process the input previous context vector and the following dialogue vector through a bidirectional encoder to obtain a following dialogue question sentence corresponding to the input text information.

[0089] Furthermore, the information matching module 20 includes:

[0090] Similarity submodule 21, used to calculate the similarity between the input text information and each original conversation data in the conversation database through a full-text search engine, and obtain a similarity matching value corresponding to each original conversation data;

[0091] The data acquisition submodule 22 is used to extract N original conversation data with similarity matching values ​​exceeding a similarity threshold from the conversation database, and obtain N target original data, where N is a positive integer;

[0092] The information determination submodule 23 is configured to determine, based on a preset probability, the preceding dialogue information that matches the input text information and the following dialogue information that corresponds to the preceding dialogue information from the N target original data.

[0093] Furthermore, the question generating device further includes:

[0094] The length identification submodule 301 is used to identify the length of the input preceding information and the following dialogue information to obtain a dialogue length result, wherein the dialogue length result is the byte length information of the input preceding information and the following dialogue information;

[0095] The length-filling submodule 302 is used to fill the lengths of the input preceding and following conversation information according to the preset conversation length when the conversation length result is less than the preset conversation length;

[0096] The length interception submodule 303 is configured to intercept the lengths of the input preceding and following dialogue information according to the preset dialogue length and preset interception rules when the dialogue length result is greater than the preset dialogue length.

[0097] Furthermore, the question generating device further includes:

[0098] a question updating module 50 for updating the number of subsequent dialogue questions corresponding to the input text information and determining whether the number of subsequent dialogue questions reaches a preset number of cycles M, where M is a positive integer;

[0099] The question loop module 60 is configured to, when the number of contextual dialogue questions does not reach a preset number of loops M, return to executing the process of extracting N original dialogue data with similarity matching values ​​exceeding a similarity threshold from the dialogue database to obtain N target original data, where N is a positive integer; determine, based on a preset probability, the previous dialogue information that matches the input text information and the next dialogue information corresponding to the previous dialogue information from the N target original data; replace the previous dialogue information with the input text information to obtain input context information; perform vector conversion on the input context information and the next dialogue information to obtain an input context vector and a next dialogue vector; and process the input context vector and the next dialogue vector using a bidirectional encoder to obtain next dialogue questions corresponding to the input text information, until M dialogue questions to be confirmed that match the input text information are obtained.

[0100] The question calculation module 70 is used to calculate the similarity of the M dialogue questions to be confirmed based on the input text information using the Jaccard coefficient, and update the following dialogue question corresponding to the input text information.

[0101] Furthermore, the question generating device further includes:

[0102] The data screening module 201 is used to obtain all the conversation data to be screened from the original database;

[0103] The database creation module 202 is configured to filter the conversation data to be filtered using a full-text search engine according to a preset filtering method to obtain original conversation data, and to create a conversation database based on the original conversation data.

[0104] The specific limitations of the question generation device can be found in the limitations of the question generation method above and will not be further elaborated here. Each module in the question generation device described above may be implemented in whole or in part via software, hardware, or a combination thereof. Each of these modules may be embedded in or independent of a processor in a computer device in hardware form, or may be stored in a computer device memory in software form, allowing the processor to call and execute the corresponding operations of each module.

[0105] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 8As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements a question generation method.

[0106] In one embodiment, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable by the processor. When the processor executes the computer program, the steps of the question generation method described in the above embodiment, such as steps S10 through S40, are implemented. Alternatively, when the processor executes the computer program, the functions of the modules / units of the question generation device described in the above embodiment, such as modules 10 through 40, are implemented. To avoid repetition, these are not further described here.

[0107] In one embodiment, a computer-readable storage medium is provided, storing a computer program. When executed by a processor, the computer program implements the question generation method described in the aforementioned method embodiment, or, when executed by a processor, implements the functions of the modules / units of the question generation device described in the aforementioned device embodiment. To avoid repetition, further description is omitted here.

[0108] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0109] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0110] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A question generation method, characterized in that: include: Get input text information; Matching the input text information from a conversation database to obtain previous conversation information matching the input text information and following conversation information corresponding to the previous conversation information; The following dialogue information refers to the information corresponding to the input text information and is the dialogue following the question sentence; Replacing the previous dialogue information with the input text information to obtain input previous dialogue information, and performing vector conversion on the input previous dialogue information and the next dialogue information to obtain an input previous dialogue vector and a next dialogue vector; Processing the input context vector and the context vector through a bidirectional encoder to obtain a context question corresponding to the input text information; The step of matching the input text information from a conversation database to obtain previous conversation information matching the input text information and following conversation information corresponding to the previous conversation information includes: The similarity matching algorithm in the Lucene engine is used to calculate the word frequency and inverse word frequency of each word in the input text information in the conversation database. The conversation data with the highest similarity to the input text information is obtained based on the word frequency and the inverse word frequency. The previous conversation information and the following conversation information corresponding to the previous conversation information are obtained from the conversation data with the highest similarity to the input text information.

2. The question generation method according to claim 1, wherein: The step of matching the input text information from a conversation database to obtain previous conversation information matching the input text information and following conversation information corresponding to the previous conversation information includes: Calculate the similarity between the input text information and each original conversation data in the conversation database through a full-text search engine, and obtain a similarity matching value corresponding to each original conversation data; Extracting N pieces of the original conversation data whose similarity matching values ​​exceed a similarity threshold from the conversation database, and obtaining N pieces of target original data, where N is a positive integer; According to a preset probability, previous conversation information matching the input text information and following conversation information corresponding to the previous conversation information are determined from the N target original data.

3. The question generation method according to claim 1, wherein: Before replacing the preceding dialogue information with the input text information to obtain preceding dialogue information, and performing vector conversion on the preceding dialogue information and the following dialogue information to obtain preceding dialogue vectors and following dialogue vectors, the question generation method further includes: Performing length identification on the input preceding information and the following conversation information to obtain a conversation length result, wherein the conversation length result is byte length information of the input preceding information and the following conversation information; If the result of the conversation length is less than the preset conversation length, the lengths of the input preceding information and the following conversation information are padded according to the preset conversation length; If the result of the conversation length is greater than the preset conversation length, the lengths of the input preceding information and the following conversation information are cut off according to the preset conversation length and a preset cutting rule.

4. The question generation method according to claim 2, wherein: After obtaining the following dialogue question corresponding to the input text information, the question generation method further includes: Updating the number of subsequent dialogue questions corresponding to the input text information, and determining whether the number of subsequent dialogue questions reaches a preset number of cycles M, where M is a positive integer; If the number of subsequent dialogue questions does not reach the preset number of loops M, the process returns to executing the process of extracting N pieces of original dialogue data whose similarity matching values ​​exceed a similarity threshold from the dialogue database to obtain N target original data, where N is a positive integer; determining, from the N target original data, previous dialogue information that matches the input text information and subsequent dialogue information corresponding to the previous dialogue information based on a preset probability; replacing the previous dialogue information with the input text information to obtain input previous dialogue information; performing vector conversion on the input previous information and the subsequent dialogue information to obtain an input previous vector and a subsequent dialogue vector; processing the input previous vector and the subsequent dialogue vector using a bidirectional encoder to obtain subsequent dialogue questions corresponding to the input text information, until M dialogue questions to be confirmed that match the input text information are obtained; Based on the input text information, similarity calculation is performed on the M dialogue questions to be confirmed using the Jaccard coefficient, and the following dialogue question corresponding to the input text information is updated.

5. The question generation method according to claim 1, wherein: Before matching the input text information from the dialogue database to obtain previous dialogue information matching the input text information and following dialogue information corresponding to the previous dialogue information, the question generation method further includes: Obtain all conversation data to be screened from the original database; The conversation data to be screened is screened according to a preset screening method through a full-text search engine to obtain original conversation data, and a conversation database is created based on the original conversation data.

6. A question generation device, characterized in that: The question generating device comprises: Information input module, used to obtain input text information; an information matching module configured to match the input text information from a conversation database to obtain preceding conversation information that matches the input text information and following conversation information corresponding to the preceding conversation information; the following conversation information being information corresponding to the input text information and being the following conversation information of the question sentence; a vector conversion module, configured to replace the preceding dialogue information with the input text information to obtain preceding input information, and to perform vector conversion on the preceding input information and the following dialogue information to obtain preceding input vectors and following dialogue vectors; a vector processing module, configured to process the input preceding text vector and the following dialogue vector using a bidirectional encoder to obtain a following dialogue question corresponding to the input text information; The information matching module is further configured to calculate the word frequency and inverse word frequency of each word in the input text information in the conversation database using a similarity matching algorithm in the Lucene engine, obtain the conversation data with the highest similarity to the input text information using the word frequency and the inverse word frequency, and obtain the preceding conversation information and the following conversation information corresponding to the preceding conversation information from the conversation data with the highest similarity to the input text information.

7. The question generating device according to claim 6, wherein: The information matching module includes: A similarity submodule is configured to calculate the similarity between the input text information and each original conversation data in the conversation database through a full-text search engine, and obtain a similarity matching value corresponding to each original conversation data; a data acquisition submodule, configured to extract N pieces of the original conversation data whose similarity matching values ​​exceed a similarity threshold from the conversation database, and acquire N pieces of target original data, where N is a positive integer; The information determination submodule is used to determine, based on a preset probability, the preceding dialogue information that matches the input text information and the following dialogue information corresponding to the preceding dialogue information from the N target original data.

8. The question generating device according to claim 6, wherein: The question generating device further includes: a length identification submodule, configured to identify the length of the input preceding information and the following conversation information to obtain a conversation length result, wherein the conversation length result is byte length information of the input preceding information and the following conversation information; a length-filling submodule, configured to fill in the lengths of the input preceding information and the following conversation information according to the preset conversation length when the conversation length result is less than the preset conversation length; The length interception submodule is used to intercept the length of the input preceding information and the following dialogue information according to the preset dialogue length and preset interception rules when the dialogue length result is greater than the preset dialogue length.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the question generation method according to any one of claims 1 to 5 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the question generation method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Method and device for establishing question generation model, and question generation method and device

    CN102737042A

  • Speech recommendation method and device in multi-round dialogue scene

    CN110008322A