A method, device, equipment and computer storage medium for question sentence expansion
By generating an extended question set, the problem of low efficiency and accuracy of question acquisition in the existing technology is solved, and efficient expansion and accurate acquisition of question sentences in the question-answer corpus is achieved.
Patent Information
- Application Number
- CN202011372430.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-30
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-11-30
AI Technical Summary
The prior art obtains question-and-answer corpus with low efficiency and accuracy, and lacks effective solutions.
By obtaining the pending question based on the received basic question, the pending question is obtained, and based on the semantic influence value and contextual association information of each first word in the pending question, an extended question set is generated to improve the expansion efficiency and accuracy of the question.
It realizes effective expansion of basic questions, improves the efficiency and accuracy of obtaining questions in question corpus, and avoids errors during manual construction.
Smart Images

Figure CN113392194B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, device, and computer storage medium for expanding questions. Background Art
[0002] In the construction of the question-and-answer field, question-and-answer corpora (including questions and answers associated with the questions) are crucial; however, in related technologies, questions and corresponding answers are often constructed manually as question-and-answer corpora, or historical questions of ordinary users are captured through logs, and the historical questions are manually annotated and corresponding answers are written before being used as question-and-answer corpora; but the questions in the corpora obtained by the former take a long time and may not conform to the query methods of ordinary users, and the accuracy of the obtained questions is low; the latter requires manual annotation of historical questions, which will seriously affect the efficiency and accuracy of obtaining questions in the question-and-answer corpus. Therefore, there is currently no effective solution to the problems of a small number of questions, low accuracy, and low efficiency in obtaining questions in the question-and-answer corpus in related technologies. Summary of the Invention
[0003] Embodiments of this application provide a method, apparatus, device, and computer storage medium for expanding questions, which are used to improve the efficiency and accuracy of obtaining questions in the question-and-answer corpus.
[0004] In the first aspect of this application, a method for expanding questions is provided, including:
[0005] Based on the received basic question, obtain a question to be processed;
[0006] Obtain the semantic influence value of each first word in the question to be processed, where the semantic influence value represents the degree of influence of each first word on the semantics of the question to be processed;
[0007] Based on the semantic influence values of each first word and the context association information of each first word, generate an extended question set corresponding to the question to be processed, where the extended question set includes extended questions whose semantic similarity to the question to be processed is greater than a first preset threshold, and the context association information represents the correlation between one word and each word belonging to the same question.
[0008] In the second aspect of this application, a device for expanding questions is provided, including:
[0009] An information receiving unit, configured to obtain a question to be processed based on the received basic question;
[0010] A word processing unit, configured to obtain the semantic influence value of each first word in the question to be processed, where the semantic influence value represents the degree of influence of each first word on the semantics of the question to be processed;
[0011] A question expansion unit is used to generate an expansion question set corresponding to the to-be-processed question based on the semantic influence values of each first word and the context association information of each first word. The expansion question set includes expansion questions with a semantic similarity greater than a first preset threshold to the to-be-processed question, where the context association information represents the correlation between a word and each word belonging to the same question.
[0012] In a possible implementation manner, when the information receiving unit is used to input the basic question by using a trained target neural network model and determine, as the to-be-processed question, a question with a semantic similarity greater than a second preset threshold to the to-be-processed question output by the target neural network model, the target neural network model is trained in the following manner:
[0013] Obtain an initial target neural network model, where the initial target neural network model includes an encoding network and a decoding network; the encoding network is used to learn to generate semantic vectors of each word in the question sample by using the question sample, and the semantic vector of a word is obtained by fusing the text vector of the word with the semantic information of the question sample; the decoding network is used to learn to generate a question with the same semantics as the question sample by using the semantic vectors of each word in the question sample.
[0014] Use the question samples in the first question sample set to adjust the encoding parameters of the encoding network.
[0015] Randomly initialize the decoding parameters of the initial decoding network.
[0016] Use the question samples in the second question sample set to further adjust the adjusted encoding parameters and the randomly initialized decoding parameters, and obtain a trained target neural network model based on the further adjusted encoding parameters and decoding parameters.
[0017] In a possible implementation manner, when there are at least two to-be-processed questions, the word processing unit is further used to: before obtaining the semantic influence values of each first word in the to-be-processed questions, determine a target to-be-processed question from the at least two to-be-processed questions.
[0018] The word processing unit is specifically used to: obtain the semantic influence values of each first word in the target to-be-processed question.
[0019] In a possible implementation, the word processing unit is specifically configured to: randomly select some or all of the to-be-processed questions from the at least two to-be-processed questions as the target to-be-processed questions; or screen out the to-be-processed questions from the at least two to-be-processed questions whose semantic similarity to the basic question is greater than a third preset threshold, and determine them as the target to-be-processed questions.
[0020] In a possible implementation, the question expansion unit is further configured to:
[0021] After generating an extended question set corresponding to the to-be-processed question based on the semantic influence values of each first word and the context association information of each first word, display at least one extended question in the extended question set on the first display page;
[0022] The question expansion unit is further configured to: in response to a question grouping instruction triggered by the first account through the second display page, obtain the to-be-grouped extended questions; and add the to-be-grouped extended questions to the question group indicated by the question grouping instruction.
[0023] In a possible implementation, the question expansion unit is further configured to:
[0024] In response to a question input operation triggered by a second account on the third display page, obtain the to-be-answered question input by the second account;
[0025] Determine the target question group containing the to-be-answered question;
[0026] Obtain the answer information associated with the target question group;
[0027] Display the answer information on the fifth display page.
[0028] In a third aspect of the present application, there is provided a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the methods described in the first aspect and any one of the possible implementation manners are implemented.
[0029] In a fourth aspect of the present application, there is provided a computer program product, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various possible implementation manners of the above first aspect.
[0030] In a fifth aspect of the present application, there is provided a computer-readable storage medium storing computer instructions, which, when run on a computer, cause the computer to execute the method as described in any one of the first aspect and any possible implementation manners thereof.
[0031] Since the embodiments of the present application adopt the above technical solutions, they have at least the following technical effects:
[0032] In the embodiments of the present application, based on the semantic influence values of each first word in the to-be-processed question sentence and the context association information of each first word, an extended question sentence set corresponding to the to-be-processed question sentence is automatically generated, realizing the expansion of the basic question sentence, saving time, and thus improving the efficiency of obtaining question sentences in the Q&A corpus and the efficiency of expanding question sentences; at the same time, it avoids generating incorrect question sentences when manually constructing the Q&A corpus, thereby improving the accuracy of obtaining question sentences in the Q&A corpus. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 FIG. is a schematic diagram of an application scenario provided by an embodiment of the present application;
[0034] Figure 2 FIG. is an example diagram of the process of a question sentence expansion method provided by an embodiment of the present application;
[0035] Figure 3 FIG. is a schematic diagram of the structure of a target neural network model provided by an embodiment of the present application;
[0036] Figure 4 FIG. is a schematic diagram of the training principle of a target neural network model provided by an embodiment of the present application;
[0037] Figure 5 FIG. is a schematic diagram of the training process of a target neural network model provided by an embodiment of the present application;
[0038] Figure 6 FIG. is a schematic diagram of the process of obtaining an extended question sentence set provided by an embodiment of the present application;
[0039] Figure 7 FIG. is a schematic diagram of the process of generating an extended question sentence provided by an embodiment of the present application;
[0040] Figure 8 FIG. is a schematic diagram of the process of generating the current word in an extended question sentence provided by an embodiment of the present application;
[0041] Figure 9 FIG. is an example diagram of a fourth display page provided by an embodiment of the present application;
[0042] Figure 10 FIG. is an example diagram of a first display page provided by an embodiment of the present application;
[0043] Figure 11 Another example diagram of the first display page provided by the embodiments of the present application;
[0044] Figure 12 An example diagram of a second display page provided by the embodiments of the present application;
[0045] Figure 13 An example diagram of a third display page provided by the embodiments of the present application;
[0046] Figure 14 A structural diagram of a neural network model system provided by the embodiments of the present application;
[0047] Figure 15 A schematic diagram of enriching the richness of extended questions provided by the embodiments of the present application;
[0048] Figure 16 Another schematic diagram of enriching the richness of extended questions provided by the embodiments of the present application;
[0049] Figure 17 A process schematic diagram of the accuracy of extended questions provided by the embodiments of the present application;
[0050] Figure 18 Another schematic diagram of the accuracy of extended questions provided by the embodiments of the present application;
[0051] Figure 19 A structural diagram of a question expansion device provided by the embodiments of the present application;
[0052] Figure 20 A structural diagram of a computer device provided by the embodiments of the present application;
[0053] Figure 21 A structural diagram of a terminal device provided by the embodiments of the present application. Detailed implementation manners
[0054] To better understand the technical solutions provided by the embodiments of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.
[0055] It should be noted that the terms "first", "second", etc. in the specification, claims and the above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0056] For the convenience of those skilled in the art to better understand the technical solutions of the present application, the following explains the basic concepts involved in the present application.
[0057] 1) Question Answering System
[0058] Question Answering System (QA): It is an advanced form of information retrieval system, which can answer questions raised by users in natural language with accurate and concise natural language; the main reason for the rise of research on question answering systems is people's need to obtain information quickly and accurately. At present, the question answering system is a research direction that has attracted much attention and has broad development prospects in the fields of artificial intelligence and natural language processing.
[0059] 2) Account, First Account and Second Account
[0060] Generally, an account is an identity identifier that can identify a user. In the embodiments of the present application, it mainly involves the account for creating question-answering corpus and the ordinary account for using the question answering system. In the embodiments of the present application, in order to distinguish the above two different accounts, the first account is used to refer to the user who creates the question-answering corpus, and the second account is used to refer to the ordinary account that uses the question answering system to answer the system; among them, the first account can be an enterprise-level user in the question answering system. The first account in the embodiments of the present application is used to input basic questions into the question answering system and receive the extended question set returned by the question answering system; the second account can input the question to be answered into the question answering system, and then the question answering system returns the answer information associated with the question to be answered.
[0061] 3) Question-answering Corpus
[0062] The question-answering corpus in the embodiments of the present application includes questions and the answer information associated with the questions.
[0063] 4) Word, First Word and Second Word
[0064] Generally, a word refers to one or more characters in a text, and the form of the word has an associated relationship with the phonetic form of the text. For example, when the language form of the text is Chinese, a word can be a Chinese character or a phrase composed of multiple Chinese characters; when the language form of the text is English, a word can be an English word or a phrase composed of multiple English words, etc.; when the language form of the text is French, Hindi, Italian, Japanese, Korean or other forms, those skilled in the art can set the corresponding form of the word according to actual needs.
[0065] In the embodiments of the present application, it mainly involves the words in each question (such as a basic question, a question to be processed or an extended question, etc.) and the words in the preset word set. For the convenience of distinction, in the embodiments of the present application, the first word is used to refer to the word in the question to be processed, and the second word is used to refer to the word in the preset word set.
[0066] 5) Semantic influence value and context association information of a word
[0067] For a certain word in a sentence, the semantic influence value of the word characterizes the degree of influence of the word on the semantics of the sentence; the context association information of the word characterizes the relevance between the word and each word in the sentence; the context association information can be but is not limited to being determined by the degree of association between the word and the word before it, and the degree of association between the word and the word after it.
[0068] 6) Bert (Bidirectional Encoder Representations from Transformer) model
[0069] The Bert model is the encoding network (Encoder) of the bidirectional Transformer; the goal of the Bert model is to train with a large-scale unlabeled corpus to obtain a semantic representation (Representation) of the text containing rich semantic information, and then fine-tune the semantic representation of the text in a specific natural language processing (Natural Language Processing, NLP) task, and finally apply it to the specific NLP task.
[0070] 7) Artificial Intelligence (AI)
[0071] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to achieve the best results. That is, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have functions of perception, reasoning, and decision-making. Artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning or deep learning.
[0072] 8) Natural Language Processing (NLP)
[0073] Natural language processing is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.
[0074] The design concept of this application will be described below.
[0075] In the process of building the Q&A field, Q&A corpora (including questions and answers associated with the questions) are crucial. However, in related technologies, questions are often constructed manually or existing questions are extended to obtain extended Q&A corpora. However, the efficiency of manually extending questions is low. A single person can only process a certain amount of data in a day. If a large number of Q&A corpora need to be created, a large number of questions need to be extended, which takes a long time. Moreover, the form of questions from ordinary users is unpredictable for those creating the Q&A corpora, resulting in low accuracy and richness of the extended questions. In related technologies, historical questions from ordinary users can also be captured through logs. If the user group is not large enough, the number of captured historical questions may be very small, and the forms of the captured historical questions tend to be single. Then, there are obvious deficiencies in the quantity and quality of the obtained Q&A corpora. In addition, the accuracy of the questions mined by the log capture method is low. For example, if the keyword is "class", the log capture method will recall some questions like "I want to listen to the song 'Class'", "Navigate me to the place called 'Class'", "What's the temperature in 'Class' today". It can be seen that the domains of these three corpora belong to music, navigation, and weather respectively, and they do not belong to the Q&A-type class domain at all. It can be seen that the accuracy of the questions captured by the log capture method is relatively low. In this case, additional technical personnel need to spend more time annotating the captured questions, which is time-consuming and laborious, and thus leads to low efficiency in obtaining questions in the Q&A corpora.
[0076] In view of this, the inventor designs a method, device, equipment, and computer storage medium for question extension. In the embodiments of this application, considering that both manually creating and log capturing questions in Q&A corpora are time-consuming and inefficient, the embodiments of this application consider obtaining questions in the Q&A corpora by extending questions, and further, more Q&A corpora can be obtained based on the extended questions. Specifically, in the embodiments of this application, a to-be-processed question is obtained based on a basic question, and an extended question set corresponding to the to-be-processed question (that is, an extended question set corresponding to the basic question) is generated based on the semantic influence values of each first word in the to-be-processed question and the context association information of each first word. The extended question set may include one or more extended questions whose semantic similarity to the to-be-processed question is greater than a first preset threshold.
[0077] It should be noted that the questions involved in the embodiments of this application can be, but are not limited to, text information or voice information, and those skilled in the art can set according to actual needs.
[0078] To understand the design concept of this application more clearly, the following is an example introduction to the application scenario of the embodiments of this application. Please refer to Figure 1, a structural schematic diagram of a question-and-answer system is provided. The system includes a terminal device 100 and a question-and-answer server 200. A question-and-answer client 110 can be installed on the terminal device 100 (such as, but not limited to, including 100-1 or 100-2 in the figure, etc.). The question-and-answer client 110 is the client of the question-and-answer system, and the question-and-answer server 200 is the server of the question-and-answer system; the question-and-answer client 110 and the question-and-answer server 200 communicate with each other.
[0079] The question-and-answer client 110 (such as, but not limited to, including 110-1 or 110-2 in the figure, etc.) can, but is not limited to, send the basic question indicated by the first account through the fourth display page to the question-and-answer server 200; it can also send the question to be answered indicated by the second account through the third display page to the question-and-answer server 200; it can also, based on the instruction of the question-and-answer server 200, display an extended question or corresponding answer information, etc. in the display page provided by the question-and-answer client 110.
[0080] The question-and-answer server 200 can, but is not limited to, obtain the basic question from the question-and-answer knowledge base 300 or receive the basic question sent by the question-and-answer client 110, and obtain the question to be processed based on the basic question. Furthermore, based on the semantic influence value and context association information of each first word in the question to be processed, it generates a set of extended questions corresponding to the question to be processed; furthermore, the question-and-answer server 200 can also send the set of extended questions corresponding to the question to be processed to the question-and-answer client 110.
[0081] As an embodiment, the question-and-answer server 200 can also receive the question to be answered sent by the question-and-answer client 110, and based on the question to be answered, return the answer information associated with the target question group including the question to be answered.
[0082] The above question-and-answer server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or multiple cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms in cloud service technology (such as, but not limited to, including the servers 200-1, 200-2, or 200-3 shown in the figure); the functions of the above question-and-answer server 200 can be implemented by one or more cloud servers, or can also be implemented by one or more cloud server clusters, etc.
[0083] The terminal device 100 in the embodiments of the present application may be a mobile terminal, a fixed terminal or a portable terminal, such as a mobile phone, a site, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / video camera, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a game device, or any combination thereof, including accessories and peripherals of these devices or any combination thereof.
[0084] Based on Figure 1 the application scenario, the following is an example description of the question expansion method involved in the embodiments of the present application; please refer to Figure 2 which shows the question expansion method designed in the embodiments of the present application and is applied to the above-mentioned question-and-answer system (i.e., the above-mentioned question-and-answer server 200, or the combination of the above-mentioned question-and-answer server 200 and the question-and-answer client 110), specifically as follows:
[0085] Step S201: Obtain the question to be processed based on the basic question.
[0086] As an embodiment, the basic question can be obtained before step S201. Among them, the basic question can be obtained from the questions in the question-and-answer corpus in the above-mentioned question-and-answer database 300. For example, one or more questions are randomly selected from the questions in the question-and-answer corpus as the basic question, or according to the historical recall times of each question in the question-and-answer corpus, the questions with historical recall times greater than the first number threshold are selected from the questions in the question-and-answer corpus as the basic question; the first number threshold is not limited, and those skilled in the art can set it according to actual needs; in the embodiments of the present application, the basic question can also be obtained in response to the question input operation triggered by the first account on the fourth display page, that is, the question input to the fourth display page is determined as the basic question; in this way, the basic question to be expanded can be obtained in different ways, improving the flexibility of obtaining the basic question.
[0087] As an embodiment, the basic question obtained in the embodiments of the present application may be one or more. In order to improve the diversity of the finally obtained extended question set, in the embodiments of the present application, some or all of the questions in the basic question can be directly determined as the question to be processed; if one basic question is obtained, the basic question can be determined as the question to be processed; if multiple basic questions are obtained, some questions can be selected from the multiple basic questions to be determined as the question to be processed, or all the multiple basic questions can be determined as the question to be processed.
[0088] Further, in the process of screening out some questions from multiple basic questions as questions to be processed, some questions can be randomly screened out from multiple basic questions, or different methods can be used to screen out some questions according to the different sources of multiple basic questions; if multiple questions are obtained from the questions in the Q&A corpus, some questions can be screened out according to the number of times each basic question has been recalled historically, and the basic questions with the number of times recalled historically greater than the second number threshold can be screened out and determined as questions to be processed; if multiple basic questions are obtained in response to the above question input operation, the K basic questions with the lowest semantic similarity to other basic questions can be screened out as questions to be processed, or the K basic questions with the earliest input order can be screened out as questions to be processed, etc. Among them, the second number threshold is not limited, and those skilled in the art can set it according to actual needs, and K is a positive integer.
[0089] In the embodiments of the present application, the trained target neural network model can also be used to obtain the first similar questions of the above basic questions as questions to be processed. The above target neural network model can but is not limited to include the Bert model, Ernie model, Albert model, etc.; in the embodiments of the present application, the above basic questions (some or all of the questions in the basic questions) and the first similar questions can also be determined as questions to be processed; the first similar questions of one basic question can include one or more questions, and the first similar questions are questions with a semantic similarity greater than the third preset threshold to the above basic questions; the acquisition method of the above target neural network model will be described in detail below; in the embodiments of the present application, questions to be processed can be obtained in various ways, which improves the flexibility of obtaining questions to be processed and increases the diversity of the obtained questions to be processed.
[0090] Step S202: Obtain the semantic influence values of each first word in the question to be processed, where the semantic influence value represents the degree of influence of each first word on the semantics of the question to be processed.
[0091] As an embodiment, the first reference value and the second reference value of each first word can be obtained based on the second word in the preset word set; and the first reference value and the second reference value are normalized to determine the semantic influence value of each first word; the preset word set can be a pre-created word set, and the preset word set can but is not limited to include words with relatively high usage frequencies, etc. Those skilled in the art can set the preset word set according to actual needs;
[0092] The above first reference value represents the probability of using the first word to generate the corresponding word in the extended question, that is, the first reference value characterizes the copying probability of directly copying the first word in the question to be processed when generating the extended question; the above second reference value represents the probability of using the second word to generate the corresponding word in the extended question, and the second word is a word in the preset word set whose semantic similarity to the first word is greater than the second preset threshold, that is, the second reference value characterizes the generation probability of using the second word to generate the first word when generating the extended question; the semantic influence value obtained in the above process not only characterizes the influence degree of each first word on the semantics of the question to be processed, but also the first reference value and the second reference value involved can determine whether the words in the question to be processed can be directly copied when generating the extended question, so as to improve the efficiency and accuracy of generating the extended question.
[0093] As an embodiment, in the embodiment of the present application, the question feature vector of the semantics of the question to be processed and the word feature vectors of each first word can be obtained, and then based on the distance between the word feature vectors of each first word and the question feature vector, the semantic influence value of each first word can be determined; for example, but not limited to, directly determining the distance between the word feature vectors of each first word and the question feature vector as the semantic influence value of each first word, or after weighting the above distance, obtaining the semantic influence value of each first word; in the embodiment of the present application, the copy probability of directly copying the first word and the generation probability of using the words in the preset vocabulary to generate the first word can also be obtained through a copy mechanism (Copy mechanism, Copy) and the like during the process of generating the extended question, and then after normalizing the copy probability and the generation probability of each first word, the semantic influence value of each first word is obtained, where the Copy mechanism refers to a mechanism that locates a certain segment in the input sequence and then copies the segment to the output sequence. The detailed content of obtaining the semantic influence value of each first word using the Copy mechanism will be described below; in the above process, based on business requirements, the semantic influence value of each first word can be determined in different ways, improving the flexibility of obtaining the semantic influence value of each first word.
[0094] Step S203: Based on the semantic influence values of each first word and the context association information of each first word, generate an extended question set corresponding to the question to be processed, where the extended question set contains extended questions whose semantic similarity to the question to be processed is greater than the first preset threshold.
[0095] Specifically, in the embodiments of the present application, each first word is mapped into a word vector, and the word vector is input into a context processing neural network. The context processing neural network performs operations on the input word vector, and then obtains the context association information of each first word. The above context processing neural network can be but is not limited to LSTM, BiSTM, etc.
[0096] As an embodiment, in the embodiments of the present application, the context association information of a word characterizes the correlation between the above word and each word belonging to the same question; the context association information of a word (i.e., the current word) in a question can be but is not limited to being determined by the first association degree and the second association degree of the current word; the first association degree characterizes the association degree between the current word and the previous word of the current word in the question, and the second association degree characterizes the association degree between the current word and the next word of the current word in the question; the above first association degree can be but is not limited to the probability of the current word appearing after the previous word, and the second association degree can be but is not limited to the probability of the next word appearing after the current word. For the convenience of understanding, an example of context association information is given here: Suppose the current word is "bright", the previous word of "bright" is "float", and the next word is "female". Assume that the probability of "bright" appearing after "float" is 0.8, and the probability of "female" appearing after "bright" is 0.5. Then the context association information of "bright" can be but is not limited to 0.8×0.5 = 0.01; the above context association information of "bright" is only for illustrative purposes, and those skilled in the art can also obtain the context association information of the current word through other means, such as based on the distance between the word vector mapped by the current word and the word vector mapped by the previous word, and the sum of the distances between the word vector mapped by the current word and the word vector mapped by the next word, to determine the context association information of the current word, etc.
[0097] As an embodiment, in the embodiments of the present application, it is considered to generate multiple extended question groups in the form of grouping, and the set of extended questions in the multiple extended question groups is determined as the above extended question set. At least one extended question is included in one extended question group. The first words of different extended questions in the same extended question group can be the same, and the first words of extended questions in different extended question groups can be different; in this way, multiple extended question groups with different first words can be obtained, greatly improving the number and richness of extended questions.
[0098] As an embodiment, in the above step S203, the extended question set can be obtained through but is not limited to the DBS (Diversity Beam Search) decoding mechanism, and those skilled in the art can also flexibly set other decoding mechanisms to implement the above step S203.
[0099] As an example, the target neural network model involved in step S201 is described in detail below: The above target neural network model can be, but is not limited to, a Question Generation (QG) model; in the field of natural language processing, the QG model refers to given a piece of text and a corresponding answer, and generating a question (interrogative sentence) corresponding to the answer based on these two pieces of information.
[0100] The above target neural network model can also be an architecture composed of multiple neural networks with text processing functions. Please refer to Figure 3 , an architecture of a target neural network model is given in an embodiment of the present application. The target neural network model includes an encoding network and a decoding network; the encoding network is used for interrogative sentences and learns to generate semantic vectors of each word in the interrogative sentence. The semantic vector of a word is obtained by fusing the text vector of the above word with the semantic information of the interrogative sentence; the decoding network is used to utilize the semantic vectors of each word in the interrogative sentence and learn to generate an interrogative sentence with the same semantics as the above interrogative sentence; wherein the above encoding network and decoding network can be composed of a Convolutional Neural Network (CNN) or a Recurrent Neural Network (RNN); the above encoding network and decoding network can also be, but are not limited to, composed of Transformer units; the encoding network can be, but is not limited to, a Bert model; wherein the Bert model is a bidirectional language model, and each word can utilize the context information of the word at the same time; wherein "bidirectional" means that when the model processes a certain word, it can utilize the information of both the words before and after the word at the same time; different from the traditional language model, the Bert model does not predict the most likely current word under the condition of giving all the previous words, but randomly masks some words and makes predictions using all the unmasked words.
[0101] In the following content of the embodiment of the present application, taking the encoding network and decoding network of the target neural network model composed of Transform units as an example, the training process of obtaining the target neural network model is described; please refer to Figure 4 , in the embodiment of the present application, it can be, but is not limited to, first pre-training the encoding network, and then fine-tuning the encoding parameters of the pre-trained encoding network and the decoding parameters in the decoding network through a fine-tuning mechanism to obtain the trained target neural network model; wherein in the embodiment of the present application, it can be, but is not limited to, obtaining the encoding parameters of the encoding network through pre-training with the Bert model, and then fine-tuning the encoding parameters and decoding parameters through Fine-tune; please refer to Figure 5 , the training process of the target neural network model specifically includes the following steps:
[0102] Step S501: Obtain an initial target neural network model, where the initial target neural network model includes an encoding network and a decoding network. The encoding network is used to learn to generate semantic vectors of each word in the question sample using the question sample. The semantic vector of a word is obtained by fusing the text vector of the word with the semantic information of the question sample. The decoding network is used to learn to generate a question with the same semantics as the question sample using the semantic vectors of each word in the question sample.
[0103] Step S502: Use the question samples in the first question sample set to adjust the encoding parameters of the encoding network.
[0104] As an embodiment, the question samples in the embodiments of the present application may include word feature vector samples mapped by each word in the question sample and a question feature vector sample mapped by the question sample. In step S502, the question sample may be input into the encoding network to obtain the predicted word feature vectors and predicted question feature vectors of each word output by the encoding network. Based on the deviation between the predicted word feature vectors of each word and the corresponding word feature vector samples, and the deviation between the predicted question feature vector and the question feature vector sample, determine the first prediction deviation of the encoding network. Through the loss function of the encoding network, adjust the encoding parameters of the encoding network in the direction of reducing the first prediction deviation until the pre-training end condition of the encoding network is satisfied. When adjusting the encoding parameters, it may include, but is not limited to, gradient adjustment of the encoding parameters, etc. The pre-training end condition of the encoding network may include, but is not limited to: the training duration reaches the first duration threshold, the number of times of adjusting the encoding parameters reaches the first adjustment number threshold, or the first prediction deviation is less than the first prediction deviation threshold, etc.
[0105] The encoding network after adjusting the encoding parameters by the above method can improve the accuracy of the encoding network in encoding each word in the question to obtain word feature vectors, and improve the accuracy of the encoding network in encoding the question to obtain question feature vectors. Those skilled in the art can also adjust the encoding parameters of the encoding network in other ways, which will not be elaborated here.
[0106] Step S503: Randomly initialize the decoding parameters of the initial decoding network.
[0107] Step S504: Use the question samples in the second question sample set to further adjust the adjusted encoding parameters and the randomly initialized decoding parameters, and obtain a trained target neural network model based on the further adjusted encoding parameters and decoding parameters.
[0108] As an example, the question samples in the second question sample set may include question input samples and similar question samples whose semantic similarity to the question input samples is greater than a third preset threshold. In step S504, the question input samples can be input into a neural network model composed of an encoding network and a decoding network. Based on the second prediction deviation between the predicted similar questions output by the neural network model and the corresponding similar question samples above, the adjusted encoding parameters and the decoding parameters after random initialization are adjusted in the direction of reducing the second prediction deviation in the manner of gradient descent until the training end condition is met, and a trained target neural network model is obtained. The above training end condition may but is not limited to including: the training duration reaches a second duration threshold, the number of times of adjusting the encoding parameters or the decoding parameters reaches a second adjustment number threshold, or the second prediction deviation is less than a second prediction deviation threshold, etc.
[0109] Through the above steps S501 to S504, first, the encoding network is pre-trained, then the decoding network is randomized, and then the encoding parameters of the encoding network and the decoding parameters of the decoding network are adjusted through Fine-tune. During the process of training the target neural network model, the accuracy and depth of the adjusted parameters are increased, and thus the accuracy of the first similar questions generated by the trained target neural network model is improved. For example, if a question is "You are so cute", the first similar question generated by the target neural network model before training may be "Ahhhh love ahhh", but the first similar questions generated by the above-trained target neural network can be "You are really cute" or "You look very cute", etc.
[0110] As an example, the following further describes the first reference value, the second reference value, and the process of obtaining the semantic influence values of each first word involved in step S202.
[0111] In the embodiments of the present application, the first reference value and the second reference value for obtaining each first word can be flexibly set; if a first word is only included in the to-be-processed question but not included in the preset word set, and there is no second word in the preset word set whose semantic similarity to the first word is greater than the second threshold, then the first reference value of the first word can be set to 1, and the second reference value of the first word can be set to 0; if a first word is included in the to-be-processed question and included in the preset word set, then the first reference value of the first word can be set to 0, and the second reference value of the first word can be set to 1; if a first word is included in the to-be-processed question and not included in the preset word set, but there is a second word in the preset word set whose semantic similarity to the first word is greater than the second threshold, then based on the semantic similarity between the second word and the first word, the first reference value of the first word can be set between 0 and 0.5, and the second reference value of the first word can be set between 0.5 and 1, etc.
[0112] As an embodiment, in the embodiments of the present application, the first reference value and the second reference value of each first word can also be generated through the Copy mechanism, and then the first reference value and the second reference value are normalized to obtain the semantic influence value of each first word; Copy mechanism: It is a decision layer after the decoding network, used to determine whether each word is directly copied from the original word or a new word is generated. There are two modes in the Copy mechanism when generating words. One is the word generation mode, and the other is the word copy mode. The generation model is a probability model that combines the two modes, where the first reference value is the probability of the copy mode, and the second reference value is the probability of the generation mode; adopting the Copy mechanism in the embodiments of the present application can solve the problem of out-of-vocabulary words (OOV, words not included in the above preset word set), that is, it can directly copy the out-of-vocabulary words in the to-be-processed question when generating the extended question, so as to improve the semantic similarity between the extended question library and the to-be-processed question.
[0113] As an embodiment, there is no excessive limitation on the method of normalizing the first reference value and the second reference value in step S202 above, and those skilled in the art can set it according to actual needs. For example, it can be but not limited to normalizing the first reference value and the second reference value based on the principles of the following formula (1), formula (2), or formula (3).
[0114] P i =P i _1+P i _2 Formula (1)
[0115] P i =P i _1×k1+P i _2×k2 Formula (2)
[0116]
[0117] In the above formulas (1) to (3), P i represents the semantic influence value of the i-th (i is a positive integer) first word in the question to be processed, P i _1 is the first reference value of the i-th first word above, P i _2 is the second reference value of the i-th first word above, k1 is the first weight value of the first reference value, k2 is the second weight value of the second reference value, k3 is the third weight value of the first reference value, and k4 is the weight value of the second reference value; among them, the setting methods of the above k1 to k4 are not overly limited, and those skilled in the art can set them according to actual needs.
[0118] As an embodiment, after obtaining the first reference value and the second reference value of each first word through the Copy mechanism, the first reference value and the second reference value can be processed, but not limited to, through the principle of the following formula (4) to obtain the semantic influence value of each first word.
[0119] P(y t |s t ,y t-1 ,c t ,M) = P(y t ,c|s t ,y t-1 ,c t ,M) + P(y t ,g|s t ,y t-1 ,c t ,M) Formula (4)
[0120] In formula (4), M is the set of input hidden layer states in the Copy mechanism, t represents the word at time t, c t is the attention score, s t is the hidden state of the source; g represents the probability of generating the corresponding word in the extended question using the second word, and c represents the probability of generating the corresponding word in the extended question using the first word; c|s t represents the first reference value of the word at time t, and g|s t represents the second reference value of the word at time t.
[0121] In the related art, when the basic sentence is expanded directly by the QG model, the sentence expanded by the QG model that is similar to the basic question sentence may not be smooth, and the semantics of the expanded sentence may be inconsistent with the semantics of the basic sentence; for example, the basic sentence before expansion is "How to watch the replay of XX class", then the sentence expanded by the QG model is "How to view the replay in XX class", while the sentence expanded after adopting the above-mentioned Copy mechanism is "How to view the replay in XX class". Obviously, it is not clear whether "XX class" in "How to view the replay in XX class" refers to a house-like hall or a classroom; furthermore, the Copy mechanism can well solve the problem of uncommon characters, such as the original sentence is "There is The sentence expanded by the QG model is "The arrival of famous teachers makes Tencent Classroom [unk] [unk] shine", where [unk] [unk] are the characters corresponding to the uncommon Chinese character "绮绮", but people who see the sentence do not know the semantics of [unk] [unk]. However, the sentence expanded by the Copy mechanism is "The arrival of famous teachers makes Tencent Classroom shine". Obviously, the application of the Copy mechanism can improve the accuracy of the expanded sentence. Therefore, in the embodiment of the present application, the Copy mechanism is used to determine the semantic influence value of each first word in the question to be processed, so as to improve the accuracy of the expanded question obtained when the question to be processed is expanded.
[0122] As an embodiment, the method of expanding the question to be processed in the form of groups in step S203 is further described below.
[0123] See also Figure 6 , you can obtain the extended question set corresponding to a question to be processed by the following steps:
[0124] Step S601: determining a threshold N (N is a positive integer) of the number of extended questions in an extended question set.
[0125] As an embodiment, the quantity threshold N may be preset, and those skilled in the art may set N according to actual needs, such as but not limited to setting N to 3, 5, 6 or 9.
[0126] Step S602: based on the semantic influence values of the first words in the question to be processed, first words corresponding to the largest top N semantic influence values are selected.
[0127] Step S603: determine the screened N first words as the first words of the expanded question in each of the N expanded question groups.
[0128] As an embodiment, when the semantic influence value of a first word is greater than or equal to the semantic influence threshold, it indicates that when generating an extended question sentence, the first word can be directly copied. Therefore, when the semantic influence value of the selected first word is greater than or equal to the semantic influence threshold, the first word can be directly used as the first word in the corresponding extended question sentence group; when the semantic influence value of the selected first word is less than the semantic influence threshold, a second word in the preset word set whose semantic similarity to the first word is greater than the second preset threshold can be used as the first word in the corresponding extended question sentence group; and the first words of the extended question sentences in different extended question sentence groups are different. The extended question sentences in the obtained multiple extended question sentence groups are more diverse in words under the condition of relatively similar semantics. Therefore, the richness of the extended question sentences in the obtained multiple extended question sentence groups is improved.
[0129] For ease of understanding, a specific example is given here. Please refer to Figure 7 , assuming that the question to be processed is "How many months of unemployment benefits can I receive", N is 3, and the 3 first words with the largest semantic influence values are "I", "unemployment", and "can". The semantic influence values of "I" and "unemployment" are greater than the semantic influence threshold, and the semantic influence value of "can" is less than the semantic influence threshold. Then, "I" and "unemployment" can be used as the first words of the extended question sentences in the first extended question sentence group and the second extended question sentence group respectively, and "may", which has a semantic similarity greater than the second preset threshold to "can" in the preset word set, is determined as the first word of the extended question sentence in the third extended question sentence group.
[0130] Step S604, for each extended question sentence group in the N extended question sentence groups, according to the above-mentioned context association information of the first word and the context association information of the first words other than the above-mentioned first word, obtain the extended question sentences in each extended question sentence group.
[0131] For ease of understanding, a schematic example is given here. The context association information of the current word in the extended question sentence is determined by the first association degree and the second association degree of the current word; the first association degree represents the probability of the current word appearing after the previous word; the second association degree represents the probability of the next word appearing after the current word; please refer to Figure 8 , for example, for Figure 7For the first set of extended questions in , the first word of the extended question is "I". When generating the second word of the extended question (i.e., the second word is the current word above), assuming that the two words "can" and "will" may appear after "I", and the probability of "can" appearing after "I" is 0.9 (i.e., the first correlation degree of "can" is 0.9), the probability of "lead" appearing after "can" is 0.5 (i.e., the second correlation degree of "will" is 0.5), the probability of "will" appearing after "I" is 0.1 (i.e., the first correlation degree of "can" is 0.1), and the probability of "lead" appearing after "will" is 0.9. Then the probability of "I can" appearing (a form of the above context correlation information) is 0.9×0.5 = 0.45, and the probability of "I will" appearing is 0.1×0.9 = 0.01. Since 0.01 is much smaller than 0.45, the word that appears after "I" in this set of extended questions is "can"; but after generating "industry", the probability of "industry insurance" appearing is 0.36, and the probability of "industry fund" appearing is 0.35. The values of 0.35 and 0.36 are relatively close. Therefore, after "industry", both "insurance" and "fund" are generated; the above content is only an exemplary explanation for easy understanding and does not constitute a limitation on the question expansion method provided by the embodiments of the present application.
[0132] Step S605, use the extended questions in each of the above extended question sets to generate an extended question set corresponding to the to-be-processed question.
[0133] Specifically, some or all of the extended questions can be screened out from each of the above extended question sets, and the set composed of the screened extended questions is determined as the extended question set corresponding to the to-be-processed question; for example, the set composed of all the extended questions in each extended question set can be determined as the extended question set corresponding to the to-be-processed question, or one extended question can be screened out from each of the above extended question sets to form the extended question set corresponding to the to-be-processed question, etc. Those skilled in the art can flexibly set it according to business requirements.
[0134] As an embodiment, please refer to Figure 9 , the embodiments of the present application provide an example of a fourth display page. The first account can, but is not limited to, trigger a question input operation through the fourth display page to indicate a basic question. The first account can also input answer information related to the basic question through the fourth display page. After the first account logs in to the fourth display page 900, the avatar and name of the first account can be displayed at the account display area 901 in the fourth display page 900; the first account can input a basic question in the first information input box 903 in the question and answer management area 902, and input answer information corresponding to the basic question in the second information input box 904, etc.; the second account can also search for corresponding questions or answer information through the information search box 905.
[0135] As an example, after the above step S203, the extended questions in the extended question set can also be displayed to the first account, and the extended questions can be grouped based on the instructions of the first account. Specifically, at least one extended question in the above extended question set can be displayed on the first display page. Then, in response to the question grouping instruction triggered by the first account through the second display page, the extended questions to be grouped are obtained, and the above extended questions to be grouped are added to the question group indicated by the above question grouping instruction. The second display page can be a page embedded in the first display page, or the second display page can also be a page independent of the first display page.
[0136] For ease of understanding, please refer to Figure 10 , here is an example of the first display page. In the first display page 1000, the basic question is "How are the navigation lights of an airplane distributed?" The extended questions of the questions to be processed obtained from the basic question (such as, but not limited to, including where the navigation lights of the airplane are installed, how the navigation lights on the airplane are installed, or where the navigation lights are distributed on the airplane as shown in the figure) can be displayed in the first display area 1101, and the answer information corresponding to the basic question can be displayed in the second display area 1102.
[0137] Please refer to Figure 11 and Figure 12 , here is another example of the first display page. In the first display page 1100, when testing similar questions (i.e., the extended questions), the information of the data to be fused for the basic question (i.e., 1 piece of data to be fused as shown in the figure, and this 1 piece of data to be fused is the information of the extended question extended from "What can you answer") can be displayed. The first account can enter the second display page 1200 after clicking on the information of the data to be fused.
[0138] The first account can trigger the question grouping instruction through the first control 1201 on the second display page 1200, or can also cancel the grouping operation through the second control, and can confirm the above question grouping instruction or the above grouping operation cancellation instruction through the confirmation control 1203, or cancel the above question grouping instruction or the above grouping operation cancellation instruction through the cancellation control 1204, etc.
[0139] As an example, in the embodiments of the present application, the obtained question group and the answer information associated with the question group can be saved in the Q&A database. Then, when the first account uses the Q&A system, the corresponding answer information can be returned to the first account based on the answer information of the question group in the Q&A database. Specifically, in response to the question input operation triggered by the first account on the third display page, the to-be-answered question input by the first account can be obtained; the question group containing the to-be-answered question can be determined; the answer information associated with the determined question group can be obtained; and the answer information can be displayed on the fifth display page. The fifth display page can be a page embedded in the third display page, or a page independent of the third display page, etc.
[0140] Please refer to Figure 13 , an example of the third display page is provided here. In the third display page 1300, the second account can input the to-be-answered question in the question input box 1301 and request the corresponding answer information from the Q&A system through the search control 1302. Then, the Q&A system will display the answer information associated with the target question group containing the to-be-answered question in the answer display area 1304 of the fifth display page 1303. The target question group is the question group in the Q&A knowledge base. Then, the second account can export the answer information of the to-be-answered question from the Q&A system through the answer export control 1305. The second account can also use the error feedback control 1306 to feedback that the displayed answer information is incorrect, etc.
[0141] As an example, the to-be-processed question obtained in step S201 may be one, or may include at least two. Therefore, when the to-be-processed question includes at least two, for each to-be-processed question, through step S202 and step S203, the extended question set corresponding to each to-be-processed question can be obtained; or before step S202, the target to-be-processed question can be screened out from at least two to-be-processed questions, and then in step S202, for each target to-be-processed question, the semantic influence value of each first word in the target to-be-processed question can be obtained.
[0142] Further, in order to improve the accuracy of the obtained extended question set, the target to-be-processed question can be determined from at least two to-be-processed questions by, but not limited to, the following methods:
[0143] The first question screening method: Randomly select some or all of the to-be-processed questions from the at least two to-be-processed questions as the target to-be-processed questions.
[0144] The second question screening method: Screen out the to-be-processed questions from the at least two to-be-processed questions whose semantic similarity to the basic question is greater than the third preset threshold, and determine them as the target to-be-processed questions.
[0145] The following provides a specific example of question expansion. In this example, the neural network model system is composed as shown in Figure 14 the neural network model system includes an encoding network, a decoding network, a Copy mechanism, and a DBS decoding mechanism; where:
[0146] The encoding network is used to receive a basic question, perform encoding processing on the basic question, generate semantic vectors of each word in the basic question, and transmit the generated semantic vectors of each word to the decoding network;
[0147] The decoding network performs decoding processing on the semantic vectors of each word in the basic question, obtains a to-be-processed question whose semantic similarity to the basic question is greater than a third preset threshold, and transmits the to-be-processed question to the Copy mechanism;
[0148] The Copy mechanism obtains a first reference value and a second reference value of each first word in the to-be-processed question based on a second word in a preset word set, and performs normalization processing on the obtained first reference value and second reference value to determine the semantic influence value of each first word;
[0149] The DBS decoding mechanism is used to determine the quantity threshold N in the set of expansion questions to be generated, and based on the semantic influence values of each first word and the context association information of each first word, generate N groups of expansion questions corresponding to the to-be-processed question, and generate a set of expansion questions corresponding to the to-be-processed question (i.e., the set of expansion questions corresponding to the basic question) based on the expansion questions in the N groups of expansion questions. The specific method can refer to the above content and will not be repeated here.
[0150] Please refer to Table 1. Here, a comparison of the effects of generating expansion questions of the basic question using the Beam Search (BS) decoding mechanism and generating expansion questions of the basic question using the method provided in the embodiments of the present application is given.
[0151] Table 1: Comparison of the effects of expanding questions in different ways
[0152]
[0153] It can be clearly seen from Table 1 that the character similarities of the expansion questions 1-5 expanded by the BS decoding mechanism are very high, and the richness of the expanded expansion questions is low; while the character similarities of the expansion questions 1-5 expanded by the technical solution provided in the embodiments of the present application are low, and the richness of the expanded expansion questions is high.
[0154] As described above, we can observe that there are differences in the extended questions expanded in different ways. In the embodiments of the present application, the diversity metric is used to specifically measure the richness of the expanded questions (by counting the number of distinct ngrams in a set of diversity beam search results, dividing by the total number of words in this set of beams, and then taking the average of the results for all data). Please refer to Figure 15 and Figure 16 , where the abscissa in the figure represents the number of characters and the ordinate represents the diversity metric. It can be seen that whether in the social security field or the game field, the richness of the expanded questions obtained by expanding the basic questions through the DBS decoding mechanism in the embodiments of the present application is significantly higher.
[0155] Please refer to Figure 17 and Figure 18 , which respectively provide comparison charts of the experimental results of the semantic accuracy of the expanded questions obtained by using different methods to expand the basic questions in the social security field and the medical field. It can be seen that when the method provided in the embodiments of the present application is not used, the overall accuracy of the expanded questions is very low. After using the method provided in the embodiments of the present application to perform generalization metrics during the question expansion process, the accuracy of the expanded questions in different fields (such as but not limited to including the social security field and the game field shown) has been significantly improved.
[0156] In summary, in the embodiments of the present application, based on the semantic influence values of each first word in the to-be-processed question and the context association information of each first word, an extended question set corresponding to the to-be-processed question is automatically generated, saving time and thus improving the efficiency of expanding the question; and since the situation of generating incorrect questions when manually expanding the question is avoided, the semantic accuracy of the expanded questions is also improved; in addition, the to-be-processed questions are grouped for expansion, and the first words of the expanded questions in different extended question groups are different, thereby further improving the richness of the expanded questions.
[0157] Please refer to Figure 19 , based on the same inventive concept, the embodiments of the present application provide a question expansion device 1900, including:
[0158] An information receiving unit 1901, configured to obtain a to-be-processed question based on the received basic question;
[0159] A word processing unit 1902, configured to obtain the semantic influence value of each first word in the to-be-processed question, where the semantic influence value represents the influence degree of each first word on the semantics of the to-be-processed question;
[0160] A question expansion unit 1903 is configured to generate an expansion question set corresponding to the to-be-processed question based on the semantic influence values of each first word and the context association information of each first word. The expansion question set includes expansion questions with a semantic similarity greater than a first preset threshold to the to-be-processed question, where the context association information represents the correlation between one word and each word belonging to the same question.
[0161] As an embodiment, the word processing unit 1902 is specifically configured to:
[0162] Obtain a first reference value and a second reference value for each first word based on a second word in a preset word set; the first reference value represents the probability of using the first word to generate a corresponding word in the expansion question, and the second reference value represents the probability of using the second word to generate a corresponding word in the expansion question. The second word is a word in the preset word set with a semantic similarity greater than a second preset threshold to the first word; perform normalization processing on the first reference value and the second reference value to determine the semantic influence value of each first word.
[0163] As an embodiment, the question expansion unit 1903 is specifically configured to: determine a quantity threshold N for the expansion questions in the expansion question set; screen out the first words corresponding to the top N largest semantic influence values based on the magnitudes of the semantic influence values of each first word; respectively determine the first words of the expansion questions in each of the N expansion question groups as the first words of the expansion questions in each of the N expansion question groups; for each of the N expansion question groups, obtain the expansion questions in each expansion question group according to the context association information of the first word and the context association information of the first words other than the first word; use the expansion questions in each of the expansion question groups to generate an expansion question set corresponding to the to-be-processed question.
[0164] As an embodiment, the information receiving unit 1901 is specifically configured to: determine some or all of the questions in the basic question as the to-be-processed question; or use a trained target neural network model, input the basic question, and determine the question with a semantic similarity greater than a third preset threshold to the basic question output by the target neural network model as the to-be-processed question.
[0165] As an embodiment, when the information receiving unit 1901 is configured to use a trained target neural network model, input the basic question, and determine the question with a semantic similarity greater than a second preset threshold to the to-be-processed question output by the target neural network model as the to-be-processed question, the target neural network model is trained in the following manner:
[0166] Obtain an initial target neural network model, where the above initial target neural network model includes an encoding network and a decoding network; the above encoding network is used to use the question samples to learn and generate semantic vectors of each word in the question samples, and the semantic vector of a word is obtained by fusing the text vector of the above word with the semantic information of the question samples; the above decoding network is used to use the semantic vectors of each word in the question samples to learn and generate a question with the same semantics as the question samples.
[0167] Use the question samples in the first question sample set to adjust the encoding parameters of the above encoding network.
[0168] Randomly initialize the decoding parameters of the above initial decoding network.
[0169] Use the question samples in the second question sample set to further adjust the adjusted encoding parameters and the randomly initialized decoding parameters, and obtain a trained target neural network model based on the further adjusted encoding parameters and decoding parameters.
[0170] As an embodiment, the above question to be processed includes at least two, and the above word processing unit 1902 is further used for: before obtaining the semantic influence values of each first word in the question to be processed, determine a target question to be processed from the above at least two questions to be processed.
[0171] The word processing unit 1902 is specifically used for: obtaining the semantic influence values of each first word in the target question to be processed.
[0172] As an embodiment, the word processing unit 1902 is specifically used for: randomly selecting some or all of the questions to be processed from the above at least two questions to be processed as the above target question to be processed; or screening out the questions to be processed with a semantic similarity greater than a third preset threshold to the above basic question from the above at least two questions to be processed, and determining them as the above target question to be processed.
[0173] As an embodiment, the question expansion unit 1903 is further used for: after generating an extended question set corresponding to the above question to be processed based on the semantic influence values of each first word and the context association information of each first word, display at least one extended question in the extended question set on the first display page.
[0174] The question expansion unit 1903 is further used for: in response to a question grouping instruction triggered by the above first account through the second display page, obtaining the questions to be grouped for expansion; adding the above questions to be grouped for expansion to the question group indicated by the above question grouping instruction.
[0175] As an embodiment, the question expansion unit 1903 is further configured to: in response to a question input operation triggered by the second account on the third display page, obtain the question to be answered input by the second account; determine a target question group containing the question to be answered; obtain answer information associated with the target question group; and display the answer information on the fifth display page.
[0176] As an embodiment, Figure 19 the device in can be used to implement any one of the question expansion methods discussed above.
[0177] The above-mentioned generation device 1900, as an example of a hardware entity, Figure 20 is a computer device as shown. The computer device includes a processor 2001, a storage medium 2002, and at least one external communication interface 2003; the processor 2001, the storage medium 2002, and the external communication interface 2003 are all connected through a bus 2004.
[0178] A computer program is stored in the storage medium 2002;
[0179] When the processor 2001 executes the computer program, it implements a generation method for an intelligent contract for testing blockchain services discussed above.
[0180] Figure 20 In, a processor 2001 is taken as an example, but actually the number of processors 2001 is not limited.
[0181] Among them, the storage medium 2002 may be a volatile memory, such as a random-access memory (RAM); the storage medium 2002 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), or the storage medium 2002 is any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The storage medium 2002 may be a combination of the above storage media.
[0182] Based on the same inventive concept, an embodiment of the present application provides a terminal device 100, which will be introduced below.
[0183] Please refer to Figure 21, the terminal device 100 includes a display unit 2140, a processor 2180, and a memory 2120. Among them, the display unit 2140 includes a display panel 2141, which is used to display information input by the user or information provided to the user, as well as various operation interfaces and display pages of the Q&A client 110, etc. In the embodiments of the present application, it is mainly used to display the interfaces of the clients installed in the terminal device 100, shortcut windows, etc.
[0184] Optionally, the display panel 2141 can be configured in the form of a liquid crystal display (LCD) or an organic light-emitting diode (OLED), etc.
[0185] The processor 2180 is used to read a computer program and then execute the method defined by the computer program. For example, the processor 2180 reads the application of the Q&A client, etc., so as to run the application on the terminal device 100 and display the interface of the application on the display unit 2140. The processor 2180 may include one or more general-purpose processors, and may also include one or more DSPs (Digital Signal Processors) for performing related operations to implement the technical solutions provided in the embodiments of the present application.
[0186] The memory 2120 generally includes internal memory and external memory. The internal memory can be a random access memory (RAM), a read-only memory (ROM), a cache (CACHE), etc. The external memory can be a hard disk, an optical disc, a USB flash drive, a floppy disk, or a tape drive, etc. The memory 2120 is used to store computer programs and other data. The computer program includes application programs corresponding to the clients, etc. Other data may include the operating system or data generated after the application programs are run. This data includes system data (such as configuration parameters of the operating system) and user data. In the embodiments of the present application, program instructions are stored in the memory 2120, and the processor 2180 executes the program instructions in the memory 2120 to implement any one of the question expansion methods described in the previous figures.
[0187] In addition, the terminal device 100 may further include a display unit 2140, which is configured to receive input digital information, word information, or contact touch operations or non-contact gestures, and generate signal inputs related to user settings and function controls of the terminal device 100. Specifically, in the embodiments of the present application, the display unit 2140 may include a display panel 2141. The display panel 2141, such as a touch screen, can collect touch operations of the user on or near it (such as the user using a finger, a stylus, or any suitable object or accessory to operate on the display panel 2141 or near the display panel 2141), and drive corresponding connection devices according to a preset program. Optionally, the display panel 2141 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the player and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into touch point coordinates, and then sends it to the processor 2180, and can receive and execute commands sent by the processor 2180. In the embodiments of the present application, if the user clicks on the Q&A client 110, the touch detection device in the display panel 2141 detects the touch operation, and then sends the signal corresponding to the detected touch operation to the touch controller. The touch controller converts the signal into touch point coordinates and sends them to the processor 2180. The processor 2180 determines that the user needs to operate on the Q&A client 110 according to the received touch point coordinates.
[0188] Among them, the display panel 2141 can be implemented in multiple types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the display unit 2140, the terminal device 100 may further include an input unit 2130. The input unit 2130 may include, but is not limited to, an image input device 2131 and other input devices 2132. The other input devices 2132 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.
[0189] In addition to the above, the terminal device 100 may further include a power supply 2190 for powering other modules, an audio circuit 2160, a near field communication module 2170, and an RF circuit 2110. The terminal device 100 may further include one or more sensors 2150, such as an acceleration sensor, a light sensor, a pressure sensor, etc. The audio circuit 2160 specifically includes a speaker 2161 and a microphone 2162, etc. For example, the terminal device 100 can collect the user's voice through the microphone 2162 and perform corresponding operations, etc.
[0190] As an embodiment, the number of processors 2180 may be one or more. The processors 2180 and the memory 2120 may be coupled or relatively independent.
[0191] As an embodiment, Figure 21 the processor 2180 in Figure 19 can be used to implement the functions of the information receiving unit 1901, the word processing unit 1902, and the question sentence expansion unit 1903 as in
[0192] As an embodiment, Figure 21 the processor 2180 in
[0193] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing computer program can be stored in a computer-readable storage medium. When the computer program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical discs and other various media that can store program codes.
[0194] Alternatively, if the above-mentioned integrated unit is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention essentially or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the above methods of the various embodiments of the present invention. And the foregoing storage medium includes: removable storage devices, ROM, RAM, magnetic disks, or optical discs and other various media that can store program codes.
[0195] Based on the same technical concept, an embodiment of the present application also provides a computer-readable storage medium storing computer instructions. When the above computer instructions run on a computer, the computer is caused to execute the question sentence expansion method as described above.
[0196] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0197] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.
Claims
1. A method for expanding questions, characterized in that, it includes: obtaining at least two questions to be processed based on a basic question and at least one of the questions whose semantic similarity to the basic question is greater than a third preset threshold; wherein, the question whose semantic similarity to the basic question is greater than the third preset threshold is obtained by inputting the basic question into a trained target neural network model; determining a target question to be processed from the at least two questions to be processed; obtaining the semantic influence value of each first word in the target question to be processed, where the semantic influence value represents the degree of influence of each first word on the semantics of the target question to be processed; generating an extended question set corresponding to the target question to be processed based on each first word, combining the semantic influence value of each first word and the context correlation information of each first word, where the extended question set contains extended questions whose semantic similarity to the target question to be processed is greater than a first preset threshold, and where the context correlation information represents the correlation between a word and each word belonging to the same question.
2. The method according to claim 1, characterized in that, the obtaining the semantic influence value of each first word in the target question to be processed includes: for each first word, performing the following operations: obtaining a first reference value and a second reference value of the first word based on the attribution relationship between the first word and a preset word set and a second word in the preset word set; the first reference value represents the replication probability of directly copying the first word when generating an extended question, and the second reference value represents the generation probability of generating the first word using the second word when generating an extended question, and the second word is a word in the preset word set whose semantic similarity to the first word is greater than a second preset threshold; performing a normalization process on the first reference value and the second reference value to determine the semantic influence value of the first word.
3. The method according to claim 1, characterized in that, the generating an extended question set corresponding to the question to be processed based on each first word, combining the semantic influence value of each first word and the context correlation information of each first word includes: determining a quantity threshold N of the extended questions in the extended question set, where N is a positive integer; screening out the first words corresponding to the top N largest semantic influence values based on the magnitudes of the semantic influence values of each first word; respectively determining the first words of the extended questions in each extended question group in the N extended question groups as the first words of the extended questions in each extended question group; for each extended question group in the N extended question groups, obtaining the extended questions in each extended question group according to the context correlation information of the first word and the context correlation information of the first words other than the first word; generating an extended question set corresponding to the question to be processed using the extended questions in each extended question group.
4. The method according to claim 1, characterized in that, When the basic question is input into the trained target neural network model, and the question output by the target neural network model with a semantic similarity greater than a second preset threshold to the question to be processed is determined as the question to be processed, the target neural network model is trained in the following manner: An initial target neural network model is obtained. The initial target neural network model includes an encoding network and a decoding network. The encoding network is used to learn and generate semantic vectors of each word in the question sample by using the question sample. The semantic vector of a word is obtained by fusing the text vector of the word with the semantic information of the question sample. The decoding network is used to learn and generate a question with the same semantics as the question sample by using the semantic vectors of each word in the question sample. Use the question samples in the first question sample set to adjust the encoding parameters of the encoding network. Randomly initialize the decoding parameters of the initial decoding network. Use the question samples in the second question sample set to further adjust the adjusted encoding parameters and the randomly initialized decoding parameters, and obtain the trained target neural network model based on the further adjusted encoding parameters and decoding parameters.
5. The method according to claim 1, wherein, Determining the target question to be processed from the at least two questions to be processed includes: Randomly selecting some or all of the questions to be processed from the at least two questions to be processed as the target question to be processed; or Selecting the questions to be processed with a semantic similarity greater than a third preset threshold to the basic question from the at least two questions to be processed, and determining them as the target questions to be processed.
6. The method according to any one of claims 1-5, wherein, After generating the extended question set corresponding to the question to be processed based on the semantic influence values of each first word and the context association information of each first word, it further includes: Display at least one extended question in the extended question set on the first display page; The method further includes: Responding to the question grouping instruction triggered by the first account through the second display page, and obtaining the extended questions to be grouped; Adding the extended questions to be grouped to the question group indicated by the question grouping instruction.
7. The method according to claim 6, wherein, The method further includes: Responding to the question input operation triggered by the second account on the third display page, and obtaining the question to be answered input by the second account; Determine the target question group containing the question to be answered; Obtain the answer information associated with the target question group; Display the answer information on the fifth display page.
8. A question extension device, wherein, It includes: An information receiving unit, configured to obtain at least two questions to be processed based on at least one of the received basic question and the question with a semantic similarity greater than a third preset threshold to the basic question. Among them, the question with a semantic similarity greater than a third preset threshold to the basic question is obtained by inputting the basic question into the trained target neural network model. A word processing unit, configured to determine a target question to be processed from the at least two questions to be processed, and obtain semantic influence values of each first word in the target question to be processed, where the semantic influence value represents the influence degree of each first word on the semantics of the target question to be processed; A question expansion unit, configured to generate an expansion question set corresponding to the target question to be processed based on each first word, in combination with the semantic influence value of each first word and the context correlation information of each first word, where the expansion question set includes expansion questions whose semantic similarity to the target question to be processed is greater than a first preset threshold, and the context correlation information represents the correlation between one word and each word belonging to the same question.
9. The apparatus according to claim 8, wherein, the word processing unit is specifically configured to: for each first word, perform the following operations: Based on the belonging relationship between the first word and a preset word set, and a second word in the preset word set, obtain a first reference value and a second reference value of the first word; the first reference value represents the replication probability of directly replicating the first word when generating an expansion question, and the second reference value represents the generation probability of generating the first word by using the second word in the preset word set, where the second word is a word in the preset word set whose semantic similarity to the first word is greater than a second preset threshold; Perform normalization processing on the first reference value and the second reference value to determine the semantic influence value of each first word.
10. The apparatus according to claim 8, wherein, the question expansion unit is specifically configured to: Determine a quantity threshold N of the expansion questions in the expansion question set, where N is a positive integer; Based on the magnitudes of the semantic influence values of each first word, screen out the first words corresponding to the top N largest semantic influence values; Respectively determine the screened N first words as the first words of the expansion questions in each expansion question group among the N expansion question groups; For each expansion question group among the N expansion question groups, obtain the expansion questions in each expansion question group according to the context correlation information of the first word and the context correlation information of the first words other than the first word; Generate an expansion question set corresponding to the question to be processed by using the expansion questions in each expansion question group.
11. A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the program, the steps of the method according to any one of claims 1-7 are implemented.
12. A computer-readable storage medium, wherein, the computer-readable storage medium stores computer instructions, and when the computer instructions run on a computer, the computer is caused to execute the method according to any one of claims 1-7.
Citation Information
Patent Citations
A duplication statement generation method and device
CN109710915A